Skip to content
Lecture Revision Lab

Methods

How the lab works, how it was evaluated, and where it falls short

The method behind every page, the evaluation design, the assumptions it rests on and its limits. Decisions are recorded as they were made, including the numbers that did not come out the way I hoped.

1 · Data

Data provenance

  • Demo text. Chapters 1, 2 and 12 of Statistics by Barbara Illowsky and Susan Dean (OpenStax, Rice University, CC BY 4.0), fetched from the openstax/osbooks-statistics repository at a pinned commit (7dea80a) and converted to plain text with exercises, homework, calculator notes and labs removed: 19 sections and 345 passages. Mathematics is rendered as plain text.
  • Your uploads. PDF, TXT and Markdown files are parsed, chunked and embedded in your browser and stored only in its IndexedDB. Passages retrieved from them are sent to a model provider only when you generate with your own key; in the Library's whole-document mode a passage is a whole file, and Generate and Ask say so first.
  • Gold set. 40 questions written from the demo text, each with verbatim evidence spans, and 480 pooled relevance judgements: 119 answer-bearing and 135 related passages. The file is web/src/data/eval/gold-set.json.
  • The 2024 database held private tutoring notes. It stays in the repository for the record and is never read by the app or deployed.

2 · Method

Retrieve, generate, review

  1. Chunk. Text is split into sentences and packed into passages of at most 900 characters, carrying up to 180 characters of trailing sentences into the next passage (DR-002).
  2. Embed. all-MiniLM-L6-v2, the 2024 model, runs in a Web Worker. Each passage is embedded once and cached; the demo text's vectors are precomputed.
  3. Retrieve. By default every passage is scored by cosine similarity, exactly as the 2024 retriever did. BM25 keyword ranking and a hybrid (reciprocal rank fusion of the dense and BM25 top 100, k = 60) are options on Retrieve.
  4. Generate. With your own key, the numbered passages and your topic go to the model with fixed grounding rules: use only the passages, cite them as [n], and say when they do not cover something. Output is JSON validated against a schema; citations outside the retrieved passages are removed and flagged.
  5. Review. Automatic checks are shown next to each item. You accept, edit or reject each one, and the decision is recorded against the call in the AI log. Only accepted items are exported.

3 · Evaluation

Evaluation design

Full results, per-question rankings and downloads are on Evaluation; the reasoning is in DR-003.

  • Retrieval gold set. Student-style questions across six types, verbatim evidence spans for answer-bearing passages, and a depth-10 pool of all three rankers judged blind to the ranker before any metric was computed. Unjudged passages count as not relevant. The chunking comparison scores sections, so every section either indexing unit ranks in its top 3 was also read for every question (128 question-section pairs, added after review; none held a further answer).
  • Metrics. nDCG@10 is the primary metric, declared in the analysis code before the first run (in the same session, so not independently time-stamped). MRR@10, Recall@5 and @10, and Success@1, 3, 5 and 10 are secondary.
  • Uncertainty. The unit is the question. Means get percentile bootstrap intervals (10,000 resamples, seed 20241009); hit rates get Wilson intervals. Rankers are compared on the same resampled questions, with Cohen's d_z and a sign-flip permutation test; Holm's adjustment covers the three primary comparisons. Differences in hit rates use Newcombe's paired score interval with McNemar's exact test.
  • Headline. Hybrid minus dense on nDCG@10 is +0.048 (0.008 to 0.089), Holm-adjusted p = 0.084: a positive point estimate that is not significant after Holm's adjustment.
  • Generation. A bring-your-own-key harness runs eight fixed topics through your model, applies automatic checks (citations, numbers, word overlap, answerability, multiple-choice form, distractors, duplicates) and collects your rubric ratings. Rates get Wilson intervals with the sample size divided by a design effect estimated from how much they vary between topics. The automatic checks stay hidden until you rate an item as faithful or not, and only those blind ratings enter the raw agreement and Cohen's kappa between the check and your judgement. A topic whose output is refused, cut off or invalid counts as a failure of the model, and evaluation calls do not use the provider's refusal fallback.
  • Verification. Every published retrieval metric, interval and test, including the chunking comparison and the breakdown by type, is recomputed in Python with numpy, scipy, statsmodels and pytrec_eval (dense and section rankings from the stored vectors), the paired score intervals are checked in R with contingencytables, and the report is a test snapshot. BM25 and hybrid rankings come from the app's own code, and a test checks they match what Retrieve shows.

4 · Assumptions

What the numbers rest on

  • The 40 questions stand in for the questions students ask about this text, and each question is an independent draw from that population, which is what the bootstrap and permutation tests resample.
  • A passage holding a verbatim evidence span answers the question, and a passage outside the pooled top 10 of all three rankers would not change the conclusions. The chunking comparison does not need this assumption: every section either indexing unit ranked in its top 3 was read.
  • The textbook is correct: passages are treated as ground truth, and faithfulness means agreeing with them.
  • Browser queries are embedded by the 8-bit model, while the evaluation used full-precision Python query vectors. The vectors agree to a cosine of at least 0.996 for queries and 0.991 for chunks, but agreement of the resulting rankings on the gold set has not been measured; the evaluation shows the same ranking code given the same query vector.
  • For generation, the items of one pack are correlated, so their intervals use a design effect estimated across packs (topics), and kappa's interval resamples whole packs.

5 · Limitations

Limitations

  • Model-drafted labels. Claude Opus 5.5 drafted the questions and relevance judgements under the owner's direction, and no person has reviewed them yet, so the retrieval results are provisional.
  • The questions were written while reading the evidence, which favours word overlap and so BM25.
  • Forty questions give wide intervals: a true difference of 0.03 in nDCG@10 is well within the noise.
  • One corpus, a statistics textbook. Lecture slides, with their fragments and bullet points, may behave differently.
  • No generation scores are published, because there is no budget for a reference run; the automatic checks are proxies, and their agreement with a person is measured only when someone rates a run.
  • Mathematics is plain text after conversion, which hurts both embeddings and readability.

6 · Next

What I'd change

  • Have a second person label the pools independently and report agreement on relevance.
  • Use questions written by students who have not seen the passages, or past exam questions.
  • Grow the set to about 63 questions (80% power for the observed effect, d_z = 0.36, at the 5% level) or about 84 at the first Holm threshold.
  • Compare chunk sizes and a stronger small embedder on the gold set before changing the default ranker.
  • Tighten the Content-Security-Policy to nonce-based scripts so 'unsafe-inline' can go, and add an optional per-session spending cap.
  • With a small budget, publish one labelled reference run of the generation evaluation.

7 · Transparency

AI use statement

How artificial intelligence is used in Lecture Revision Lab, what it never does, and what data leaves your browser.

What AI does here

  • Drafts revision material. With your own API key, a language model drafts a revision agenda, worked examples or likely student questions from passages retrieved from your notes or the demo textbook. Each item must cite the passages it uses.
  • Answers follow-up questions on the Ask page, from retrieved passages, with citations.
  • Takes part in the generation evaluation at /evaluation/generation, where you can run a fixed set of topics through your chosen model and rate the results.
  • Embeds text for search. A small embedding model (all-MiniLM-L6-v2) runs in your browser to turn passages and questions into vectors. It does not generate text.

What AI never does here

  • It never grades, marks or makes any decision about a person.
  • It never runs without your key, and no server receives your key: calls go from your browser straight to the provider. The key is stored only in your browser for this site (in session storage, or in local storage if you tick "remember"), where scripts running on this site could read it, and it is never written to the AI log.
  • It never produces the retrieval rankings, the retrieval evaluation or its statistics. Those are computed by deterministic code and checked against independent Python implementations.
  • It never puts an item you have not accepted into a revision export: the Markdown, CSV and Anki files contain only the items you accepted. The AI log and evaluation exports are different on purpose: they contain every output, including pending and rejected ones, so the record can be audited.
  • It never sends your other records to a provider: not the AI log, saved packs, evaluation runs or files that were not retrieved. Only the passages retrieved for a request leave the browser (listed below). With the default sentence chunks a passage is a short extract; in the Library's whole-document mode (the 2024 behaviour) a passage is an entire uploaded file, so retrieved files are sent in full, and Generate and Ask say so before you send anything.

What is sent to the provider, and when

Only when you press Generate, Ask or Run evaluation with a key set:

  • the system prompt (a persona plus fixed grounding rules);
  • the numbered passages retrieved for your topic or question: typically three to eight passages of up to 900 characters each with the default chunking, up to 2,000 characters if you raise the chunk size in the Library, or whole files (or whole textbook sections on Generate's whole-section index) in whole-document mode;
  • your topic or question;
  • on the Ask page, earlier questions and answers in the conversation, unless you switch history off.

Your key goes in the request header to the provider. Nothing is sent to this site, which has no server code that receives your data. The provider's terms and data-retention policy apply to what you send it.

People stay in the loop

  • Every generated item is a draft. You accept, edit or reject it, and only accepted items are exported.
  • Ask answers can be marked useful or not useful.
  • Each decision is recorded against the call that produced the output in the AI log.

Transparency

  • Every AI output is labelled "AI-generated", with the model that produced it. Pre-recorded examples, shown when no key is set, are labelled as recordings.
  • The AI log at /ai-log lists every call, successful or not: the feature, provider, model, input (including any earlier conversation turns sent with it), output, latency, token usage and your decision. It can be searched and exported as JSON or CSV, and cleared at any time. It lives only in your browser.
  • Exported revision material keeps the label: the Markdown export marks each item AI-generated, the CSV has ai_generated and model columns, and Anki cards carry an ai-generated tag.
  • Automatic checks on generated material (citations, numbers, word overlap, question quality) are shown next to each item, with their definitions and limits. In the generation evaluation they appear only once you have rated the item as faithful or not, so your rating is not steered by them.

How AI was used to build this site

The 2026 rebuild was written with help from Claude Opus 5.5, which also drafted the retrieval gold set's questions and relevance judgements under the owner's direction and wrote the labelled pre-recorded examples. Every evidence span in the gold set is checked verbatim against the textbook by a script. The owner has not yet reviewed the labels, so the retrieval results are marked provisional until that review is done and recorded in the gold set.

Frameworks

This statement and the design behind it are informed by the Australian Government's policy for the responsible use of AI in government, the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. The site does not claim compliance with any of them.

8 · System card

Intended use, failure modes and rights

System card: Lecture Revision LabA study aid, not a grading tool. Components and data, retrieval results with intervals, known failure modes (hallucinated citations, invented numbers, retrieval misses, prompt injection through uploads) and the rights you need in uploaded material.Read the system card

9 · Decisions

Decision records

Each record states the decision first, then the context, the options, why, what happened (weak numbers included) and what I would change. Records are never edited after the fact; a later one supersedes them.

  1. DR-001 · Accepted · decided 9 October 2026Browser-only architecture with bring-your-own-key AI and a local audit logRun the whole lab in the visitor's browser, make every AI feature optional and powered only by the visitor's own API key called directly from the browser, and record every model call with the human decision on its output in an audit log kept in that browser.
  2. DR-002 · Accepted · decided 9 October 2026Sentence-aware 900-character chunks with the 2024 MiniLM embeddingsKeep the prototype's embedding model, all-MiniLM-L6-v2, but index sentence-aware passages of at most 900 characters with up to 180 characters of overlap instead of whole documents, cache every passage vector, keep dense cosine ranking as the default, and offer BM25 and a reciprocal-rank-fusion hybrid as options on Retrieve.
  3. DR-003 · Accepted · decided 9 October 2026Evaluate retrieval on a pooled gold set with paired intervals, and generation with checks and a human rubricMeasure retrieval on a 40-question gold set with verbatim evidence and pooled, blind relevance judgements, drafted by a model under my direction and reviewed by me before the results count as final; compare the three rankers on the same questions with paired intervals and a primary metric declared before the first run; and measure generation through a bring-your-own-key harness that combines transparent automatic checks with a human rating rubric and reports how often the two agree.