Decision record DR-003 · Accepted · decided 9 October 2026
Evaluate retrieval on a pooled gold set with paired intervals, and generation with checks and a human rubric
Decision: Measure retrieval on a 40-question gold set with verbatim evidence and pooled, blind relevance judgements, drafted by a model under my direction and reviewed by me before the results count as final; compare the three rankers on the same questions with paired intervals and a primary metric declared before the first run; and measure generation through a bring-your-own-key harness that combines transparent automatic checks with a human rating rubric and reports how often the two agree.
| Status | Decided | Recorded | Owner |
|---|---|---|---|
| Accepted | 9 October 2026 | 10 October 2026 | Sunchuangyu (Rin) Huang |
Context
The revival proved parity with the 2024 code: the same rankings, the same chunk spans, near-identical vectors. Parity says nothing about whether the retriever finds the passage that answers a question, or whether the generated revision material is faithful to the passages it cites. Those are the questions a student, a tutor or a reviewer of this project would ask.
The measurement has to be honest at a small sample size, reproducible from a fixed seed, checkable by someone else, and free to run. There is no budget for model calls, so any generation measurement must run on the visitor's own key and publish nothing the site did not compute.
Decision
Retrieval:
- Authoring: Claude Opus 5.5 drafted the questions, evidence spans and relevance judgements under my direction, reading all 19 sections. I review every question and label before the results count as final, and record the date and any changed labels in the gold set (
authoring.review). Until then the site labels the retrieval results provisional. - Forty questions about the demo textbook (12 on sampling and data, 17 on descriptive statistics, 11 on regression and correlation), phrased the way a student would ask and spread over six types: definition, procedure, interpretation, comparison, conceptual and numeric. The app's ten sample queries are excluded.
- Each question lists one or more evidence spans copied verbatim from the text. Every chunk that contains a whole span is answer-bearing (grade 2). A script checks every span against the corpus.
- The top 10 chunks of each ranker were pooled and every pooled chunk was read and graded, sorted by chunk id with no indication of which ranker returned it, before any metric was computed: grade 1 if it is about the same specific point but does not answer alone, grade 0 otherwise. A pooled chunk that did answer the question got a new evidence span instead. Unjudged chunks count as not relevant.
- Primary metric, declared in the analysis code before the first run (in the same working session, so not independently time-stamped): nDCG@10 with graded gains. Secondary: MRR@10, Recall@5 and Recall@10 over answer-bearing chunks, and Success@1, 3, 5 and 10 (an answer-bearing chunk in the top k).
- The unit of analysis is the question. Means get 95% percentile bootstrap intervals (10,000 resamples, seed 20241009) and hit rates get Wilson intervals. Each pair of rankers is compared on the same resampled questions (paired bootstrap of the mean difference), with Cohen's d_z, wins, ties and losses, and a two-sided sign-flip permutation test (exact up to 16 non-zero differences, otherwise 20,000 Monte Carlo sign vectors). Holm's adjustment covers the three primary comparisons only. Differences in hit rates (Success@5, and Success@1 and @3 in the chunking comparison) get Newcombe's paired score interval and McNemar's exact test instead of the bootstrap interval.
- A second comparison holds the ranker fixed (dense) and changes only the indexing unit: whole sections, as in 2024, against chunks collapsed to their sections. Because it scores sections, it has its own pool: every section that either arm ranks in its top 3 is read for each question, so a miss is never just a section nobody read. A test fails if a question's top 3 sections in either arm are neither answer-bearing nor recorded as read.
- Every metric, interval and test on
/evaluation, including Success@3, the breakdown by question type and the whole chunking comparison, is recomputed independently in Python with numpy, scipy, statsmodels and pytrec_eval, with the dense and section rankings recomputed from the stored vectors. The BM25 and hybrid rankings come from the app's own code, and a test checks they are exactly what Retrieve shows. The published report is a test snapshot, so the code, the gold set and the page cannot drift apart.
Generation (bring your own key):
- Eight fixed topics, one per main area of the text. Each retrieves its top five chunks with the dense ranker using precomputed query vectors, and asks the visitor's chosen model for likely student questions (four multiple-choice and three short-answer). Temperature is 0 where the model accepts it, and every call is in the audit log.
- Automatic checks on every item: it cites a retrieved passage; every citation points to one; numbers in the item appear in the passages it cites; at least half of its content words do; and for questions, the answer key is in the cited passages, multiple-choice items have four distinct options and a valid key, distractors are on topic and of similar length to the key, and no question is a near-duplicate of an earlier one.
- A human rubric on each item: faithful to the cited passages, correct, answerable from them, distractors plausible, useful for revision, and an accept or reject decision.
- Blind rating: an item's automatic checks stay hidden until the rater has given the "faithful" rating. The rater can choose to see them earlier, and every rating records whether the checks were on screen first (
checksShownBeforeRating, also in the exports). Only blind ratings enter the agreement between the checks and the rater, because a rater who has seen the verdict is anchored by it. - A topic whose output the model refuses, cuts off or returns as invalid JSON counts as a failure of the model, reported as the share of topics with valid output (Wilson interval over topics). Evaluation calls do not opt in to Anthropic's server-side refusal fallback, so a refusal cannot be answered by another model and scored as the evaluated one; a pack that a fallback model did write counts as a refusal. Failures outside the model (key, rate limit, network, a stop) are left out and can be re-run. Item-level rates are among successful topics.
- Item-level rates get a Wilson interval with the sample size divided by Kish's design effect, 1 + (m − 1) × ICC, where m is the mean pack size and the intraclass correlation is estimated from how much the rate varies between topics, because the items of one pack share a prompt, passages and a model call. Agreement between the automatic support check and the human faithfulness rating is reported as raw agreement and Cohen's kappa over blind ratings; kappa's interval resamples whole topics, resamples where kappa is undefined are dropped and counted, and kappa is not reported as a measure of agreement when either rater gave every item the same verdict. Runs export as JSON or CSV.
Options considered
| Option | Why not |
|---|---|
| Judge every chunk for every question | 13,800 judgements for 40 questions; pooling the top 10 of each ranker covers what any of them could show a student. |
| Questions and labels generated by a model and used without review | Tunes the set to a generator's phrasing and to whatever it found easy, and nobody checks a single label. |
| Questions and labels written entirely by hand | About two days of one person's time before any result; a model draft reviewed label by label gets most of the benefit sooner. |
| Recall@k alone | Ignores the order of results and the difference between answering and merely relating to a question. |
| Report means without intervals | With 40 questions, differences of a few hundredths are noise; a bare mean invites over-reading. |
| A model as the judge of faithfulness | Costs money on every run and has a model grade a model; worth adding later as a third rater with its own agreement figure. |
| Model-drafted gold set reviewed by the owner, pooled judgements, paired intervals, checks plus rubric (chosen) | Small, reproducible and free, with the uncertainty visible and every number checkable; the review and the "provisional" label keep a model's judgement from passing as a person's. |
Why
Pooling is the standard way to build relevance judgements when exhaustive labelling is too expensive, and judging without knowing the ranker removes the most obvious bias. A primary metric declared before the first run and Holm's adjustment over a small family keep the number of ways to find a "significant" result small. A paired design uses each question as its own control, which matters at n = 40. For generation, simple checks with published definitions are easier to trust than an opaque score, and measuring their agreement with a person shows how far to trust them.
What happened
- The gold set has 40 questions, 119 answer-bearing chunks and 135 related chunks, from 480 pooled judgements.
- On nDCG@10 the hybrid scored 0.794 (0.745 to 0.839), BM25 0.778 (0.722 to 0.829) and dense 0.746 (0.674 to 0.811). Hybrid minus dense was +0.048 (0.008 to 0.089), d_z = 0.36, with 28 wins, 1 tie and 11 losses; the permutation p-value was 0.028, which becomes 0.084 after Holm's adjustment, so the difference is not significant at the 5% level by the rule declared before the first run. Hybrid minus BM25 (+0.016) and dense minus BM25 (−0.032) had intervals well across zero.
- BM25 was stronger than I expected: an answer-bearing chunk in the top five for all 40 questions (Wilson interval 0.91 to 1.00) and the best MRR@10, 0.827 (0.740 to 0.908). Dense missed the top five on four questions, for example "Which point does the least-squares regression line always pass through?", where the answer is a short phrase inside a long derivation.
- Chunking beat whole-section vectors (see DR-002): +0.25 in section-level Success@1 (Newcombe paired score interval 0.068 to 0.415), McNemar exact p = 0.013.
- The breakdown by question type is exploratory, with three to thirteen questions per type, and the page says so.
- No generation numbers are published, because there is no budget for a reference run. The harness, the checks and the rubric are tested with mocked calls and a pre-recorded pack.
- Weak points, stated on the page: the questions and judgements were drafted by Claude Opus 5.5 under my direction and I have not yet reviewed them, so every retrieval number is labelled provisional; questions written while reading the evidence, which favours word overlap and so BM25; the analysis plan was declared in code in the same session as the first run rather than committed beforehand; and 40 questions, which gives wide intervals.
- A second review before merge (10 October 2026) found four problems in the evaluation itself, all fixed before merge:
- The judgement pool came only from the three chunk rankers, so the 2024 whole-section ranker in the chunking comparison was judged only on sections the chunk rankers happened to reach; for five of its fourteen first-place misses the section had no judged chunk at all. Every section either arm put in its top 3 was then read for every question (128 question-section pairs). None held a further answer, so no evidence was added and the result (36 of 40 against 26 of 40) stands, but it now rests on a complete section-level pool rather than on the pooling assumption. This reading was done after the first results, not before them.
- The rubric showed each item's automatic checks above the rating controls, so the agreement between checks and rater was inflated by anchoring. Ratings are now blind by default and only blind ratings count.
- The numeric grounding check misread lists and ranges in the cited passages ("18, 19, 20" kept only 20, and the range "18-25" was read as −25), so correct items were flagged as unsupported; one number tokenizer now reads items and passages the same way.
- The topic bootstrap gave zero-width intervals (100% to 100%) whenever every topic passed a check, and under-covers with eight topics; rates now use the cluster-adjusted Wilson interval above.
- Review before merge (10 October 2026) changed three things. An earlier draft of this record rejected "questions written by a model" while doing exactly that; the options now say what was done. The paired bootstrap interval for Success@5 differences was too narrow with a handful of discordant questions all pointing one way (dense minus BM25: −0.200 to −0.025, while McNemar's exact p was 0.125); those rows now use Newcombe's paired score interval (−0.231 to 0.006). And the generation summary first dropped failed topics from every rate, so a model that often returned invalid JSON scored as well as one that never failed; model faults now count against the model.
What I'd change
- Read the section pool before the first results rather than after a review, and have a second person check those 128 judgements along with the rest.
- Commit the analysis plan (primary metric, comparison family, alpha) in its own commit before generating any results.
- Have a second person label the pools independently and report inter-annotator agreement on relevance.
- Use questions written by students who have not seen the passages, or past exam questions, to remove the lexical advantage.
- Grow the set: with the observed effect (d_z = 0.36), about 63 questions would give 80% power for one paired test at the 5% level, and about 84 at the Holm threshold for the first of three comparisons.
- Add a model-as-judge rater to the generation harness, reported only through its agreement with the human rubric.
Source: docs/decisions/DR-003-evaluation-design.md in the project repository (GenAI-Lec-Gen, private).