Personal GenAI prototype · 2024 · rebuilt 2026
Turn lecture notes into revision you can check.
Lecture Revision Lab finds the passages that answer a question, drafts revision agendas, worked examples and practice questions from them, and makes every item point back to its source. Nothing leaves the page until you have reviewed it.
- Runs in your browser
- Bring your own key, or none
- Demo text: OpenStax, CC BY 4.0
The idea
Good revision material starts with finding the right paragraph.
GenAI-Lec-Gen started in October 2024 as a small Python prototype for preparing course material: put lecture slides into a knowledge base, ask for a range of content, and get back a revision agenda, detailed examples and the questions students are likely to ask. It used retrieval-augmented generation, so answers would come from the slides rather than the model's memory.
The prototype stopped halfway. Storage and retrieval worked, but the orchestration files were empty and the generation step never ran end to end. This rebuild finishes the idea as a browser app, keeps the original retriever and data model, and adds what was missing: chunking, cached embeddings, citations, structured output and a human review step.
How it works
Five steps, each one visible
- 01
Chunk
Text is packed into passages of whole sentences, about 900 characters each, with a little overlap.
- 02
Embed
all-MiniLM-L6-v2, the 2024 model, runs in a Web Worker. Each passage is embedded once and cached.
- 03
Retrieve
Your question is scored against every passage by cosine similarity and the top k come back.
- 04
Generate
With your own key, the model drafts structured output and must cite passages as [n].
- 05
Review
Accept, edit or reject each item, then export Markdown, CSV or an Anki deck.
Checked against the original
Same retriever, now in the browser
The rebuild ports the 2024 algorithm rather than replacing it. Automated tests compare the TypeScript code with numbers produced by running the original Python.
- sample queries return the same top-10 passages as the 2024 Python retriever
- 10 / 10
- TypeScript port of CustomRM checked against NumPy, scores within 0.00001
- lowest cosine between browser and Python query embeddings
- 0.996
- 8-bit Transformers.js model vs sentence-transformers; top-3 passages identical for all 10
- each passage is embedded, then cached
- Once
- the 2024 retriever re-embedded every stored document on every query
- original database tests ported to the browser store
- 4 / 4
- same tables: knowledge_base, generated_content, conversation_history
Measured, with intervals
Does it find the right passage? Measured, not claimed.
A gold set of student-style questions, pooled relevance judgements, paired comparisons and 95% intervals, every metric and test recomputed independently in Python. Weak results are reported too.
- nDCG@10 for the hybrid ranker, the highest point estimate of three
- Retrieval, 40 questionsprovisional0.794
- 95% interval 0.745 to 0.839; not distinguishable from BM25 (+0.016, −0.040 to 0.066). Its difference from the 2024 retriever, +0.048, is not significant after Holm's adjustment (p = 0.084). Provisional: the labels were drafted by a model and no person has reviewed them yet.
- questions whose answering section ranks first, chunks against whole sections
- Indexing unitprovisional36 vs 26
- Out of 40, same model and scoring, with every section either index ranked in its top 3 read for each question. McNemar exact p = 0.013. Provisional: the labels were drafted by a model and no person has reviewed them yet.
- logged with its input, output, model and your decision
- GovernanceEvery call
- No server receives your key. Outputs are labelled AI-generated, and revision exports contain only the items you accepted.
2024 prototype and 2026 rebuild
What changed, and what did not
Passages
- 2024 prototype
- One row per uploaded file; the model reads only its first 256 word pieces
- 2026 rebuild
- Sentence chunks with overlap. Whole-document mode kept for comparison
Retrieval
- 2024 prototype
- Dot product of normalised MiniLM vectors, top k = 3, every document re-embedded per query
- 2026 rebuild
- Same scoring and model, embeddings cached in IndexedDB or precomputed
Generation
- 2024 prototype
- DSPy ChainOfThought on gpt-4o-mini; the system prompt and history were built but never sent
- 2026 rebuild
- Anthropic (default) or OpenAI with your key; prompt and history sent; JSON validated with zod; citations required
Content tasks
- 2024 prototype
- content_generator.py and main.py were empty stubs
- 2026 rebuild
- Revision agenda, worked examples, likely questions, and follow-up chat
Storage
- 2024 prototype
- SQLite file with three tables
- 2026 rebuild
- The same three tables in your browser, plus an AI audit log
Evaluation
- 2024 prototype
- None; the generation step never ran end to end
- 2026 rebuild
- A 40-question retrieval gold set with intervals and paired tests, and a bring-your-own-key generation harness with a human rubric
Interface
- 2024 prototype
- A Python script run from the terminal
- 2026 rebuild
- A web app with human review and exports
| Aspect | 2024 prototype (Python) | 2026 rebuild (browser) |
|---|---|---|
| Passages | One row per uploaded file; the model reads only its first 256 word pieces | Sentence chunks with overlap. Whole-document mode kept for comparison |
| Retrieval | Dot product of normalised MiniLM vectors, top k = 3, every document re-embedded per query | Same scoring and model, embeddings cached in IndexedDB or precomputed |
| Generation | DSPy ChainOfThought on gpt-4o-mini; the system prompt and history were built but never sent | Anthropic (default) or OpenAI with your key; prompt and history sent; JSON validated with zod; citations required |
| Content tasks | content_generator.py and main.py were empty stubs | Revision agenda, worked examples, likely questions, and follow-up chat |
| Storage | SQLite file with three tables | The same three tables in your browser, plus an AI audit log |
| Evaluation | None; the generation step never ran end to end | A 40-question retrieval gold set with intervals and paired tests, and a bring-your-own-key generation harness with a human rubric |
| Interface | A Python script run from the terminal | A web app with human review and exports |
About this project
Personal project, 2024
Built and rebuilt by Sunchuangyu (Rin) Huang. Started in October 2024 as a tool for preparing revision material; revived as a browser app in October 2026. The source lives in a private GitHub repository (GenAI-Lec-Gen).
- Original stack
- Python, DSPy 2.5, OpenAI gpt-4o-mini, sentence-transformers all-MiniLM-L6-v2, SQLite, PyPDF2
- Revived stack
- Next.js 16, TypeScript, Transformers.js in a Web Worker, IndexedDB, pdf.js, zod, Tailwind CSS, shadcn/ui; Anthropic or OpenAI with your own key
- Provenance
- The 2024 code is kept unchanged in the repository. Its database held private tutoring notes, so they are not used here; the demo reads an openly licensed textbook instead.
- Privacy
- There is no server code handling your data. Files, keys and history stay in your browser, and model calls go straight to the provider you choose.
Start with a question about box plots.
The demo textbook is loaded and the sample questions answer instantly.