Decision record DR-002 · Accepted · decided 9 October 2026
Sentence-aware 900-character chunks with the 2024 MiniLM embeddings
Decision: Keep the prototype's embedding model, all-MiniLM-L6-v2, but index sentence-aware passages of at most 900 characters with up to 180 characters of overlap instead of whole documents, cache every passage vector, keep dense cosine ranking as the default, and offer BM25 and a reciprocal-rank-fusion hybrid as options on Retrieve.
| Status | Decided | Recorded | Owner |
|---|---|---|---|
| Accepted | 9 October 2026 | 10 October 2026 | Sunchuangyu (Rin) Huang |
Context
The 2024 prototype stored each uploaded file as one row and embedded the whole file with sentence-transformers/all-MiniLM-L6-v2. That model reads at most 256 word pieces, so for a lecture handout or a textbook section everything after the first page or so was invisible to retrieval. The prototype also re-embedded every stored document on every query.
The revival has to run the embedding model in a browser for free, keep the download small, and stay comparable with the 2024 retriever so that the parity tests mean something. Generated items must cite the passages they draw on, so a passage should be short enough to check by eye and should not cut a sentence in half.
Decision
- Keep all-MiniLM-L6-v2. In the browser it runs as the 8-bit ONNX export through Transformers.js (about 23 MB, cached by the browser). The demo textbook's vectors are computed once in Python at full precision and shipped as static files.
- Split text into sentences (with an abbreviation list), then pack whole sentences greedily into passages of at most 900 characters. Each new passage starts with the trailing sentences of the previous one that fit in 180 characters. The TypeScript chunker is a line-for-line port of the Python one.
- Cache each passage vector in IndexedDB under a hash of the model, precision, length limit and text, so a passage is embedded once.
- Keep the 2024 whole-document index as a switch on Retrieve so the difference can be seen.
- Rank by the dot product of unit vectors (cosine similarity), exactly as the prototype's
CustomRMdid. After the evaluation in DR-003, add BM25 and a hybrid (reciprocal rank fusion of the dense and BM25 top 100, k = 60) as visible options, without changing the default.
Options considered
| Option | Why not |
|---|---|
| Whole documents, as in 2024 | Most of a long document is never read by the model, and a citation to a whole section is hard to check. |
| Fixed-size token windows | Cuts sentences, and so claims, in half; needs the tokenizer on the main thread. |
| No overlap between passages | A definition split across a boundary is found by neither half. |
| A larger embedding model (for example bge-small or e5) | A bigger download and a different model, which breaks the parity checks against the 2024 retriever. |
| A hosted embedding API | Costs money and sends the visitor's notes to another service. |
| Sentence-aware 900-character chunks, same model (chosen) | Most passages fit inside the 256 word-piece window, sentences stay whole, and the 2024 scoring is unchanged. |
Why
Nine hundred characters is about 196 word pieces on average for the demo text (the longest chunk is 378), so most passages fit inside the model's window, and a passage is about one short paragraph, which is easy to check against a citation. Keeping the model and the scoring means every difference from 2024 comes from the indexing unit and can be measured.
What happened
- On the 19 demo sections the TypeScript chunker reproduces all 345 Python chunk spans exactly. Browser (8-bit) and Python (32-bit) vectors agree to a cosine of at least 0.991 over all 345 chunks and 0.996 over the ten sample queries. An earlier figure of 0.993 for chunks came from a sample of 24 chunks that missed every chunk longer than the window; checking all of them showed that Transformers.js dropped the closing [SEP] when it truncated and lower-cased Σ differently from the Python tokenizer (0.980 on the long chunks). The browser worker now patches both.
- Thirty-eight of the 345 chunks (11 percent) are longer than 256 word pieces, so their last sentences are not seen by the model. They are still found by BM25.
- Measured on the 40-question gold set (DR-003), chunking helps. With the same model and scoring, the top-ranked unit came from a section that answers the question for 36 of 40 questions with chunks (0.90, Wilson 95% interval 0.77 to 0.96) against 26 of 40 with whole-section vectors (0.65, 0.50 to 0.78). The paired difference is +0.250 (Newcombe paired score 95% interval 0.068 to 0.415; McNemar exact p = 0.013). A first draft of this record quoted a percentile bootstrap interval (0.100 to 0.425), which is too narrow for a paired difference of hit rates with few discordant questions. Review then found that the relevance pool had been built from the chunk rankers only, which can penalise the whole-section index for sections nobody judged; every section either index put in its top 3 was read for every question (128 question-section pairs), none held a further answer, and the figures above did not change. Like the rest of the labels, these judgements were drafted by a model and are provisional.
- Dense ranking was not the strongest ranker of chunks. On nDCG@10 the hybrid scored 0.794 (0.745 to 0.839), BM25 0.778 (0.722 to 0.829) and dense 0.746 (0.674 to 0.811). The hybrid's gain over dense, +0.048 (0.008 to 0.089), does not survive the Holm adjustment declared before the first run (adjusted p = 0.084). BM25 put an answer-bearing chunk in the top five for all 40 questions, dense for 36.
- I therefore kept dense as the default, for parity with 2024 and because the gold-set labels were drafted by a model and I have not yet reviewed them, and made BM25 and the hybrid one click away on Retrieve. A test checks that the app's three rankers produce exactly the evaluated rankings when given the same query vector. The evaluation used full-precision query vectors from Python, while queries typed in the browser are embedded by the 8-bit model; how much that changes the rankings on the gold set has not been measured.
What I'd change
- Compare chunk sizes of 600, 900 and 1,200 characters on the gold set before trusting 900.
- Embed the 40 gold questions with the 8-bit browser model and report the top-10 overlap and the change in nDCG@10 for dense and hybrid ranking.
- Try a stronger small embedder with a similar download size, now that the evaluation can show whether it is worth breaking parity.
- Make the hybrid the default if reviewed labels, a second annotator or a larger question set confirm its gain over dense.
- Render mathematics properly; the plain-text conversion of formulas hurts both embeddings and readability.
Source: docs/decisions/DR-002-chunking-and-embedding-choices.md in the project repository (GenAI-Lec-Gen, private).