Skip to content
Lecture Revision Lab
Methods

System card · October 2026

System card: Lecture Revision Lab

A retrieval-augmented study aid that finds the passages answering a question and, with the visitor's own API key, drafts revision material that must cite them. This card covers the whole system: the retrievers, the language models it can call, the data, the evaluation and the known ways it fails. It also serves as the model card for the embedding model as used here.

FieldValue
SystemLecture Revision Lab (revival of GenAI-Lec-Gen, a personal prototype from October 2024)
VersionOctober 2026
OwnerSunchuangyu (Rin) Huang
Retrievalall-MiniLM-L6-v2 cosine (default), BM25, and a reciprocal-rank-fusion hybrid
GenerationThe visitor's own key: Claude Haiku 4.5 (default), Sonnet 5.5 or Opus 5.5; or an OpenAI model they name
Where it runsEntirely in the visitor's browser; model calls go from the browser to the chosen provider
Live sitelecture-revision-lab.vercel.app

Intended use

  • A study aid: a student or tutor searches lecture material they have the right to use, and drafts revision agendas, worked examples and practice questions that point back to the passages they came from.
  • Every generated item is a draft for a person to accept, edit or reject. Revision exports (Markdown, CSV, Anki) contain only the items a person accepted; the audit log and evaluation exports keep every output, so that they can be audited.
  • A demonstration of retrieval-augmented generation with its evaluation, audit log and human review visible.

Not intended for

  • Grading, marking or any decision about a student. The system was not designed or evaluated for assessment, and its outputs must not be used as evidence of what a student knows.
  • Answering live assessment tasks, or producing material to be submitted as a student's own work.
  • Professional advice of any kind (medical, legal, financial), or replacing the source text: the cited passages remain the authority.
  • Personal data. Uploads should not contain information about identifiable people.

Components

  • Embedding model. sentence-transformers/all-MiniLM-L6-v2 (Apache-2.0), 384 dimensions, inputs truncated at 256 word pieces, mean pooled and L2 normalised. In the browser it runs as the 8-bit ONNX export (Xenova/all-MiniLM-L6-v2) through Transformers.js. It was not fine-tuned. Browser vectors agree with full-precision Python vectors to a cosine of at least 0.991 for all 345 demo chunks and 0.996 for queries, after two tokenizer patches that match sentence-transformers' truncation and lower-casing; the evaluation used the full-precision query vectors, and how far the 8-bit queries change the rankings on the gold set has not been measured.
  • Passages. Sentence-aware chunks of at most 900 characters with up to 180 characters of overlap (DR-002).
  • Rankers. Dense cosine (the 2024 CustomRM scoring), BM25 (k1 = 1.2, b = 0.75) over stemmed, stop-word-filtered tokens, and a hybrid that fuses the dense and BM25 top 100 by reciprocal rank (k = 60). Dense is the default.
  • Language models. Called only with the visitor's key (DR-001). Prompts keep the 2024 framing (numbered context passages, then the question), add fixed grounding rules (use only the passages, cite them as [n], say when they do not cover something), and ask for JSON that is validated against a schema. The model that answered is recorded, including when a server-side fallback answered.

Data

  • Demo corpus. Chapters 1, 2 and 12 of Statistics by Barbara Illowsky and Susan Dean (OpenStax, CC BY 4.0), fetched from the openstax/osbooks-statistics repository at commit 7dea80a, converted to plain text with exercises, homework and labs removed: 19 sections, 345 chunks. Mathematics is rendered as plain text, which loses some formula structure.
  • Uploads. PDF, TXT and Markdown files are parsed in the browser and stored only in that browser's IndexedDB. When the visitor generates with a key, the retrieved passages are sent to their chosen provider: short chunks by default, but entire files when the Library is set to whole-document mode (the 2024 behaviour), which Generate and Ask point out before anything is sent.
  • Gold set. Forty questions with verbatim evidence spans and 480 pooled relevance judgements (DR-003), in web/src/data/eval/gold-set.json. Claude Opus 5.5 drafted the questions and judgements under the owner's direction; the owner has not yet reviewed them, so the retrieval results are provisional.
  • The 2024 database. The prototype's SQLite file held private tutoring notes. It stays in the repository for the record and is never read by the web app or deployed.

Evaluation

Retrieval on the 40-question gold set; 95% intervals are percentile bootstrap (10,000 resamples, seed 20241009) for means and Wilson for hit rates, and differences in hit rates use Newcombe's paired score interval. The primary metric, nDCG@10, was declared in the analysis code before the first run, in the same working session, so it is not independently time-stamped. The results are provisional until the owner reviews the model-drafted labels.

RankernDCG@1095% intervalMRR@10Success@5
Dense0.7460.674 to 0.8110.75036 of 40
BM250.7780.722 to 0.8290.82740 of 40
Hybrid0.7940.745 to 0.8390.76139 of 40
  • Hybrid minus dense on nDCG@10 is +0.048 (0.008 to 0.089); after Holm's adjustment for three comparisons it is not significant at the 5% level (adjusted p = 0.084). No other primary comparison comes close.
  • Hybrid and BM25 cannot be told apart on nDCG@10 (hybrid minus BM25 +0.016, −0.040 to 0.066), and BM25 has the higher MRR@10, so "highest point estimate" is the most that can be said for the hybrid.
  • With dense ranking, chunks put an answering section first for 36 of 40 questions against 26 of 40 for the 2024 whole-section index (McNemar exact p = 0.013). Every section either index ranked in its top 3 was read for every question, so the comparison does not rest on unjudged sections.
  • Generation is evaluated by the visitor on their own key at /evaluation/generation: the share of topics that return valid output (refusals, cut-off and invalid output count as failures, and evaluation calls do not use the refusal fallback), automatic citation, grounding and question-quality checks plus a human rubric given before the checks are shown, with Wilson intervals widened by a design effect estimated across topics. The site publishes no generation scores because it has no budget for a reference run.
  • Every published retrieval metric, interval and test, including the chunking comparison, is recomputed independently in Python (numpy, scipy, statsmodels, pytrec_eval), with the dense and section rankings recomputed from the stored vectors. Full results, per-question rankings and limitations are on /evaluation.

Known failure modes

  • Hallucinated or misattributed citations. A model can cite a passage it was shown that does not support the claim, or cite a number past the passages it was given. Out-of-range citations are removed and flagged automatically. In-range citations that do not support the claim are only partly caught by the word-overlap and number checks; the reviewer has to read the source.
  • Invented numbers. Statistics material is full of numbers. The numeric check flags numbers that do not appear in the cited passages, but cannot tell a correct derived number (say, an IQR the passage implies) from an invented one.
  • Retrieval misses. The default dense ranker missed the top five for four of the 40 gold questions, usually when the answer is a short phrase inside a long derivation. Generation from the wrong passages produces fluent, well-cited and still unhelpful material.
  • Truncation. Eleven percent of demo chunks exceed the model's 256 word-piece window; their final sentences are invisible to dense retrieval.
  • Garbled formulas. Plain-text mathematics from the conversion can read oddly in passages and in generated items.
  • Prompt injection through uploads. An uploaded document can contain instructions aimed at the model. The system prompt tells the model to use the passages only as sources, outputs are validated against a schema, and every item is reviewed before export, but an injected instruction can still change what is drafted.
  • Provider behaviour. Models can refuse, cut off at the length limit or return invalid JSON; each case is shown to the visitor and logged as a failure. Browser extensions or networks that block cross-origin requests stop calls from leaving the browser.

Rights in uploaded material

  • Upload only material you have the right to use, such as your own notes or openly licensed text. The upload panel says so.
  • Uploads stay in your browser, but passages retrieved from them are sent to your chosen provider when you generate. The provider's own terms and data-retention policy then apply.
  • Generated material is derived from the passages. Keep the attribution the source licence requires; the Markdown export lists every cited source.

Ethical considerations and governance

  • A person stays in the loop: every item is accepted, edited or rejected, and that decision is written to the audit log against the call that produced it.
  • Every AI output is labelled "AI-generated"; pre-recorded examples are labelled as recordings.
  • The audit log (/ai-log) lists every call with its input (including any earlier conversation turns sent with it), output, model, latency, token usage and decision, and exports as JSON or CSV. The key is never written to it.
  • The design is informed by the Australian Government's policy for the responsible use of AI in government, the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It does not claim compliance with any of them.
  • The AI use statement on /methods says what AI does here, what it never does and what data leaves the browser.

Maintenance

Decision records live in docs/decisions and are never edited after the fact; a later record supersedes an earlier one. The gold set, the evaluation report and the Python reference values are regenerated by the scripts in scripts/ and checked by the test suite.

Source: docs/system-card.md in the project repository (GenAI-Lec-Gen, private).