Skip to content
Lecture Revision Lab

Personal GenAI prototype · 2024 · rebuilt 2026

Turn lecture notes into revision you can check.

Lecture Revision Lab finds the passages that answer a question, drafts revision agendas, worked examples and practice questions from them, and makes every item point back to its source. Nothing leaves the page until you have reviewed it.

  • Runs in your browser
  • Bring your own key, or none
  • Demo text: OpenStax, CC BY 4.0

The idea

Good revision material starts with finding the right paragraph.

GenAI-Lec-Gen started in October 2024 as a small Python prototype for preparing course material: put lecture slides into a knowledge base, ask for a range of content, and get back a revision agenda, detailed examples and the questions students are likely to ask. It used retrieval-augmented generation, so answers would come from the slides rather than the model's memory.

The prototype stopped halfway. Storage and retrieval worked, but the orchestration files were empty and the generation step never ran end to end. This rebuild finishes the idea as a browser app, keeps the original retriever and data model, and adds what was missing: chunking, cached embeddings, citations, structured output and a human review step.

How it works

Five steps, each one visible

  1. 01

    Chunk

    Text is packed into passages of whole sentences, about 900 characters each, with a little overlap.

  2. 02

    Embed

    all-MiniLM-L6-v2, the 2024 model, runs in a Web Worker. Each passage is embedded once and cached.

  3. 03

    Retrieve

    Your question is scored against every passage by cosine similarity and the top k come back.

  4. 04

    Generate

    With your own key, the model drafts structured output and must cite passages as [n].

  5. 05

    Review

    Accept, edit or reject each item, then export Markdown, CSV or an Anki deck.

Checked against the original

Same retriever, now in the browser

The rebuild ports the 2024 algorithm rather than replacing it. Automated tests compare the TypeScript code with numbers produced by running the original Python.

sample queries return the same top-10 passages as the 2024 Python retriever
10 / 10
TypeScript port of CustomRM checked against NumPy, scores within 0.00001
lowest cosine between browser and Python query embeddings
0.996
8-bit Transformers.js model vs sentence-transformers; top-3 passages identical for all 10
each passage is embedded, then cached
Once
the 2024 retriever re-embedded every stored document on every query
original database tests ported to the browser store
4 / 4
same tables: knowledge_base, generated_content, conversation_history

Measured, with intervals

Does it find the right passage? Measured, not claimed.

A gold set of student-style questions, pooled relevance judgements, paired comparisons and 95% intervals, every metric and test recomputed independently in Python. Weak results are reported too.

nDCG@10 for the hybrid ranker, the highest point estimate of three
Retrieval, 40 questionsprovisional0.794
95% interval 0.745 to 0.839; not distinguishable from BM25 (+0.016, −0.040 to 0.066). Its difference from the 2024 retriever, +0.048, is not significant after Holm's adjustment (p = 0.084). Provisional: the labels were drafted by a model and no person has reviewed them yet.
questions whose answering section ranks first, chunks against whole sections
Indexing unitprovisional36 vs 26
Out of 40, same model and scoring, with every section either index ranked in its top 3 read for each question. McNemar exact p = 0.013. Provisional: the labels were drafted by a model and no person has reviewed them yet.
logged with its input, output, model and your decision
GovernanceEvery call
No server receives your key. Outputs are labelled AI-generated, and revision exports contain only the items you accepted.

2024 prototype and 2026 rebuild

What changed, and what did not

About this project

Personal project, 2024

Built and rebuilt by Sunchuangyu (Rin) Huang. Started in October 2024 as a tool for preparing revision material; revived as a browser app in October 2026. The source lives in a private GitHub repository (GenAI-Lec-Gen).

Original stack
Python, DSPy 2.5, OpenAI gpt-4o-mini, sentence-transformers all-MiniLM-L6-v2, SQLite, PyPDF2
Revived stack
Next.js 16, TypeScript, Transformers.js in a Web Worker, IndexedDB, pdf.js, zod, Tailwind CSS, shadcn/ui; Anthropic or OpenAI with your own key
Provenance
The 2024 code is kept unchanged in the repository. Its database held private tutoring notes, so they are not used here; the demo reads an openly licensed textbook instead.
Privacy
There is no server code handling your data. Files, keys and history stay in your browser, and model calls go straight to the provider you choose.

Start with a question about box plots.

The demo textbook is loaded and the sample questions answer instantly.

Open the retrieval explorer