Skip to content
Lecture Revision Lab

Showcase · recorded with Playwright

Guided tour

Three short walkthroughs of the main workflows, and a screenshot of every key feature. A Playwright script records them from a production build of this site with fixed inputs, so they can be re-made whenever the site changes. Every video has captions and a written transcript.

Workflows

Watch the main journeys

Each video is muted, with the current step in a caption bar at the bottom of the frame. The same steps are listed beside it, and the transcript gives the time of each one.

1:05 · 2.7 MB · MP4 Download the video

Walkthrough 1

Build a knowledge base

Load the open-textbook demo corpus, see how its text becomes passages, then search it with the 2024 retriever, your own question and two other rankers. No key needed.

  1. 1Open the Library. The demo knowledge base is an open textbook: OpenStax Statistics, CC BY 4.0
  2. 219 sections become 345 passages, embedded once with the 2024 model, all-MiniLM-L6-v2
  3. 3Chunking preview: whole sentences packed into passages of about 900 characters, with overlap
  4. 4Retrieve, no key needed: a sample question returns the top passages instantly
  5. 5Each passage shows its section, cosine score and the query words it matches
  6. 6Ask your own question. The model downloads once and embeds it in your browser
  7. 7Switch to the hybrid ranker and widen k: dense and BM25 fused by reciprocal rank
Transcript

Silent screen recording of the “Build a knowledge base” workflow in Lecture Revision Lab, captioned step by step. A highlighted circle shows the pointer.

  1. 0:00Open the Library. The demo knowledge base is an open textbook: OpenStax Statistics, CC BY 4.0
  2. 0:0319 sections become 345 passages, embedded once with the 2024 model, all-MiniLM-L6-v2
  3. 0:10Chunking preview: whole sentences packed into passages of about 900 characters, with overlap
  4. 0:20Retrieve, no key needed: a sample question returns the top passages instantly
  5. 0:26Each passage shows its section, cosine score and the query words it matches
  6. 0:33Ask your own question. The model downloads once and embeds it in your browser
  7. 0:51Switch to the hybrid ranker and widen k: dense and BM25 fused by reciprocal rank
Try it yourself

1:05 · 2.4 MB · MP4 Download the video

Walkthrough 2

Generate revision questions

Add your own key, generate practice questions that cite the retrieved passages, accept, edit or reject each one, export the accepted set, and find the call in the AI log.

Mocked AI response for illustration. The AI part of this recording uses a placeholder key, and the recording script answers the call with the site's pre-recorded questions under the model name mock-claude-haiku-for-illustration. No real model was called.

  1. 1Open Generate and choose likely student questions on box plots and quartiles
  2. 2Bring your own key: Anthropic by default with Claude Haiku 4.5, kept in this tab only
  3. 3A placeholder key and a mocked AI response for illustration: no real model is called
  4. 4Generate: five passages are retrieved and every question must cite them as [n]
  5. 5Each item is labelled AI-generated and checked automatically: citations, answer key, duplicates
  6. 6Review every item: accept it, edit it, or reject it
  7. 7Export only the accepted items as Markdown, CSV or an Anki deck
  8. 8The call is in the AI log with its input, output, tokens and your decision, never the key
Transcript

Silent screen recording of the “Generate revision questions” workflow in Lecture Revision Lab, captioned step by step. A highlighted circle shows the pointer.

  1. 0:00Open Generate and choose likely student questions on box plots and quartiles
  2. 0:05Bring your own key: Anthropic by default with Claude Haiku 4.5, kept in this tab only
  3. 0:13A placeholder key and a mocked AI response for illustration: no real model is called
  4. 0:18Generate: five passages are retrieved and every question must cite them as [n]
  5. 0:25Each item is labelled AI-generated and checked automatically: citations, answer key, duplicates
  6. 0:32Review every item: accept it, edit it, or reject it
  7. 0:46Export only the accepted items as Markdown, CSV or an Anki deck
  8. 0:53The call is in the AI log with its input, output, tokens and your decision, never the key
Try it yourself

0:47 · 2.4 MB · MP4 Download the video

Walkthrough 3

How good is retrieval?

Recall@k, nDCG@10 and Success@k for three rankers on a 40-question gold set, each with a 95% interval, paired comparisons, a browser re-run of the whole report, and the AI audit log.

Mocked AI response for illustration. The AI part of this recording uses a placeholder key, and the recording script answers the call with the site's pre-recorded questions under the model name mock-claude-haiku-for-illustration. No real model was called.

  1. 1Open Evaluation: 40 gold-set questions, three rankers, labels still provisional
  2. 2The design: pooled, blind relevance judgements before any score was computed
  3. 3nDCG@10, MRR, Recall@5 and Success@5 by ranker, each with a 95% interval
  4. 4Paired differences on the same questions, with Holm-adjusted p-values and effect sizes
  5. 5Chunks against whole sections: the answering section ranks first more often (McNemar exact)
  6. 6Re-run the whole evaluation in your browser and compare it with the published report
  7. 7The AI audit log: every model call with its input, output, latency, tokens and human decision
  8. 8Inspect any record as JSON, or export the whole log as JSON or CSV
Transcript

Silent screen recording of the “How good is retrieval?” workflow in Lecture Revision Lab, captioned step by step. A highlighted circle shows the pointer.

  1. 0:00Open Evaluation: 40 gold-set questions, three rankers, labels still provisional
  2. 0:04The design: pooled, blind relevance judgements before any score was computed
  3. 0:08nDCG@10, MRR, Recall@5 and Success@5 by ranker, each with a 95% interval
  4. 0:15Paired differences on the same questions, with Holm-adjusted p-values and effect sizes
  5. 0:20Chunks against whole sections: the answering section ranks first more often (McNemar exact)
  6. 0:24Re-run the whole evaluation in your browser and compare it with the published report
  7. 0:29The AI audit log: every model call with its input, output, latency, tokens and human decision
  8. 0:35Inspect any record as JSON, or export the whole log as JSON or CSV
Try it yourself

Key features

Screenshots

Desktop screenshots are 1440 × 900. Phone screenshots are 390 × 844, rendered at twice the resolution and scaled down. Select one to see it full size.

On a phone

Reproducible

How these were made

One script

web/e2e/showcase.spec.ts drives Google Chrome through each journey with Playwright and doubles as an end-to-end test. Run it with pnpm showcase, and point BASE_URL at any deployment.

Fixed inputs

The demo textbook, the sample question, the typed question and the box-plot topic are fixed in the script, and the evaluation numbers come from the published report, so a re-recording shows the same results.

No real key

The AI steps type a placeholder key. The script answers every request to Anthropic or OpenAI inside the browser with the site's pre-recorded questions, so nothing is sent to a provider, the model is named mock-claude-haiku-for-illustration, and every frame with a mocked reply carries the label "Mocked AI response for illustration".