Evaluation
Does it find the right passage, and can you trust what the model writes?
Headline results
Provisional: the questions and relevance labels were drafted by Claude Opus 5.5 under the owner's direction and have not yet been reviewed by a person. The numbers will be recomputed after that review.
Highest nDCG@10 (point estimate)provisional
Hybrid 0.794
95% interval 0.745 to 0.839. Not distinguishable from BM25 (0.778; difference +0.016, −0.040 to 0.066); dense 0.746.
Hybrid minus the 2024 retrieverprovisional
+0.048
Paired interval 0.008 to 0.089; Holm-adjusted p = 0.084, so not significant by the rule declared before the first run.
BM25 answer in the top 5provisional
40 of 40
Wilson interval 0.91 to 1.00; dense 36 of 40, hybrid 39 of 40.
Chunks over whole sectionsprovisional
+0.25
Answering section ranked first, 36 vs 26 of 40; paired score interval 0.07 to 0.42; McNemar exact p = 0.013.
How it was measured
A pooled gold set, judged before any score was computed
- Questions
- 40, in six types, written as a student would ask; the app's sample queries are excluded.
- Relevance
- 119 answer-bearing passages (hold a verbatim evidence span) and 135 related ones.
- Pooling
- Top 10 of all three rankers, 480 judgements drafted by Claude Opus 5.5 blind to the ranker, not yet reviewed by a person. Unjudged passages count as not relevant. For the chunking comparison, every section in either arm's top 3 was read too (128 question-section pairs).
- Statistics
- Bootstrap 10,000 resamples, seed 20241009; Wilson for hit rates; sign-flip test (20,000 sign vectors when not exact); Holm over the three primary comparisons.
- MiniLM cosine (the 2024 retriever)all-MiniLM-L6-v2 vectors and a dot product of unit vectors: the prototype's CustomRM scoring, unchanged.
- BM25 keyword searchOkapi BM25 (k1 = 1.2, b = 0.75) over lower-cased, stop-word-filtered, lightly stemmed tokens: the app's offline fallback.
- Hybrid (reciprocal rank fusion)Reciprocal rank fusion of the dense and BM25 top 100 with k = 60, the value from the original paper.
Results by ranker
nDCG@10 point estimates: Hybrid 0.794, BM25 0.778, Dense 0.746
- Dense (circle)
- BM25 (square)
- Hybrid (diamond)
| Metric | Dense | BM25 | Hybrid |
|---|---|---|---|
| nDCG@10primary | 0.746 (0.674 to 0.811) | 0.778 (0.722 to 0.829) | 0.794 (0.745 to 0.839) |
| MRR@10 | 0.750 (0.642 to 0.854) | 0.827 (0.740 to 0.908) | 0.761 (0.664 to 0.853) |
| Recall@5 | 0.677 (0.567 to 0.784) | 0.729 (0.655 to 0.802) | 0.755 (0.671 to 0.833) |
| Recall@10 | 0.820 (0.722 to 0.906) | 0.856 (0.789 to 0.918) | 0.877 (0.810 to 0.935) |
| Success@1 | 25 / 40 (0.47 to 0.76) | 28 / 40 (0.55 to 0.82) | 24 / 40 (0.45 to 0.74) |
| Success@3 | 34 / 40 (0.71 to 0.93) | 38 / 40 (0.83 to 0.99) | 37 / 40 (0.80 to 0.97) |
| Success@5 | 36 / 40 (0.77 to 0.96) | 40 / 40 (0.91 to 1.00) | 39 / 40 (0.87 to 1.00) |
| Success@10 | 38 / 40 (0.83 to 0.99) | 40 / 40 (0.91 to 1.00) | 40 / 40 (0.91 to 1.00) |
Paired comparisons
Hybrid minus dense: +0.048, not significant after Holm's adjustment
| Metric | A − B | Difference (95% interval) | d_z | Wins / ties / losses | p (sign-flip) | Holm p | McNemar p |
|---|---|---|---|---|---|---|---|
| nDCG@10primary | Hybrid − Dense | +0.048 (0.008 to 0.089)bootstrap | 0.36 | 28 / 1 / 11 | 0.028MC | 0.084 | – |
| nDCG@10primary | Hybrid − BM25 | +0.016 (−0.040 to 0.066)bootstrap | 0.09 | 20 / 1 / 19 | 0.559MC | 0.968 | – |
| nDCG@10primary | Dense − BM25 | −0.032 (−0.121 to 0.050)bootstrap | −0.11 | 15 / 0 / 25 | 0.484MC | 0.968 | – |
| MRR@10 | Hybrid − Dense | +0.010 (−0.060 to 0.080)bootstrap | 0.05 | 8 / 28 / 4 | 0.792exact | – | – |
| MRR@10 | Hybrid − BM25 | −0.066 (−0.157 to 0.021)bootstrap | −0.23 | 4 / 26 / 10 | 0.162exact | – | – |
| MRR@10 | Dense − BM25 | −0.076 (−0.199 to 0.041)bootstrap | −0.19 | 7 / 21 / 12 | 0.244MC | – | – |
| Success@5 | Hybrid − Dense | +0.075 (−0.033 to 0.203)score | 0.28 | 3 / 37 / 0 | 0.250exact | – | 0.250 |
| Success@5 | Hybrid − BM25 | −0.025 (−0.129 to 0.065)score | −0.16 | 0 / 39 / 1 | 1.000exact | – | 1.000 |
| Success@5 | Dense − BM25 | −0.100 (−0.231 to 0.006)score | −0.33 | 0 / 36 / 4 | 0.125exact | – | 0.125 |
Secondary rows are descriptive: their p-values are not adjusted and do not count as findings. Ties are questions where both rankers scored the same, which is common for hit-based metrics. Differences in means use the paired bootstrap; Success@5 differences use Newcombe's paired score interval, because with a handful of discordant questions all pointing one way the bootstrap interval is too narrow (for dense − BM25 it was −0.200 to −0.025, excluding zero, while McNemar's exact p was 0.125).
Indexing unit
Chunks find the answering section far more often than whole sections
Exploratory
By question type
| Type | n | Dense nDCG@10 · S@5 | BM25 nDCG@10 · S@5 | Hybrid nDCG@10 · S@5 |
|---|---|---|---|---|
| Definition | 5 | 0.59 (0.24 to 0.91) · 3/5 | 0.86 (0.76 to 0.95) · 5/5 | 0.75 (0.56 to 0.91) · 5/5 |
| Procedure | 9 | 0.83 (0.72 to 0.92) · 9/9 | 0.80 (0.74 to 0.87) · 9/9 | 0.86 (0.82 to 0.90) · 9/9 |
| Interpretation | 6 | 0.77 (0.57 to 0.92) · 6/6 | 0.70 (0.54 to 0.86) · 6/6 | 0.80 (0.68 to 0.90) · 6/6 |
| Comparison | 4 | 0.73 (0.57 to 0.88) · 4/4 | 0.83 (0.73 to 0.93) · 4/4 | 0.82 (0.71 to 0.94) · 4/4 |
| Conceptual (why) | 13 | 0.73 (0.64 to 0.81) · 12/13 | 0.76 (0.65 to 0.86) · 13/13 | 0.75 (0.66 to 0.84) · 12/13 |
| Numeric fact | 3 | 0.81 (0.71 to 0.98) · 2/3 | 0.75 (0.40 to 0.95) · 3/3 | 0.81 (0.58 to 1.00) · 3/3 |
Every question
Per-question rankings
| Question and each ranker's top 10 |
|---|
q01 · 1.1 · ComparisonI surveyed 100 students to estimate the average for every student at the school. Which number is the parameter and which is the statistic?6 answer-bearing and 4 related passages. Evidence:
Dense 2 BM25 2 Hybrid 2 |
q02 · 1.1 · Conceptual (why)Does it make sense to calculate an average for a categorical variable such as party affiliation?4 answer-bearing and 3 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q03 · 1.2 · ProcedureHow can I tell whether quantitative data are discrete or continuous?4 answer-bearing and 6 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q04 · 1.2 · ComparisonWhen should I use a bar graph rather than a pie chart for categorical data?4 answer-bearing and 5 related passages. Evidence:
Dense 4 BM25 1 Hybrid 1 |
q05 · 1.2 · DefinitionIf I randomly pick a few whole classrooms and survey every student in them, what kind of sample is that?3 answer-bearing and 2 related passages. Evidence:
Dense 12 BM25 1 Hybrid 4 |
q06 · 1.2 · ComparisonWhat is the difference between sampling error and nonsampling error?3 answer-bearing and 3 related passages. Evidence:
Dense 2 BM25 1 Hybrid 2 |
q07 · 1.2 · Conceptual (why)Why are results from online polls that anyone can answer often unreliable?4 answer-bearing and 2 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q08 · 1.3 · Conceptual (why)Why is temperature in degrees Celsius measured on an interval scale and not a ratio scale?2 answer-bearing and 3 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q09 · 1.3 · ProcedureHow do I fill in the cumulative relative frequency column of a frequency table?4 answer-bearing and 5 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q10 · 1.4 · Conceptual (why)Why does randomly assigning subjects to treatment groups let an experiment show cause and effect?6 answer-bearing and 2 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q11 · 1.4 · Conceptual (why)Why do experiments give the control group a placebo?3 answer-bearing and 5 related passages. Evidence:
Dense 2 BM25 1 Hybrid 2 |
q12 · 1.4 · Conceptual (why)When would a researcher run an observational study instead of an experiment?1 answer-bearing and 3 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q13 · 2.1 · Numeric factIn a stem-and-leaf plot, what are the stem and the leaf of the number 432?1 answer-bearing and 2 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q14 · 2.2 · ProcedureHow do I decide how many bars a histogram should have and how wide each bar should be?3 answer-bearing and 2 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q15 · 2.2 · ProcedureIf a value lands exactly on the boundary between two histogram classes, which class does it belong to?1 answer-bearing and 3 related passages. Evidence:
Dense 2 BM25 1 Hybrid 2 |
q16 · 2.3 · InterpretationIf I scored in the 90th percentile, did I get 90 percent of the questions right?2 answer-bearing and 6 related passages. Evidence:
Dense 1 BM25 3 Hybrid 1 |
q17 · 2.3 · ProcedureHow far from the quartiles does a value have to be before it counts as a potential outlier?2 answer-bearing and 4 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q18 · 2.4 · DefinitionWhat do the whiskers of a box plot show?3 answer-bearing and 3 related passages. Evidence:
Dense 1 BM25 2 Hybrid 1 |
q19 · 2.4 · InterpretationWhen comparing two box plots, what does it mean if one box is much wider than the other?1 answer-bearing and 5 related passages. Evidence:
Dense 3 BM25 1 Hybrid 2 |
q20 · 2.5 · InterpretationIn a town where one person earns millions and everyone else earns about the same, should I report the mean or the median income?3 answer-bearing and 3 related passages. Evidence:
Dense 1 BM25 2 Hybrid 1 |
q21 · 2.5 · ProcedureHow can I estimate the mean from a grouped frequency table when the individual values are unknown?5 answer-bearing and 3 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q22 · 2.5 · Numeric factWith 100 data values, where in the ordered list is the median?2 answer-bearing and 2 related passages. Evidence:
Dense 2 BM25 5 Hybrid 3 |
q23 · 2.5 · Conceptual (why)What happens to the sample mean as I take bigger and bigger samples?2 answer-bearing and 6 related passages. Evidence:
Dense 1 BM25 3 Hybrid 3 |
q24 · 2.6 · InterpretationIf the mean of my data is bigger than the median, which way is the distribution skewed?3 answer-bearing and 4 related passages. Evidence:
Dense 1 BM25 2 Hybrid 2 |
q25 · 2.7 · Conceptual (why)Why do we divide by n minus 1 when calculating a sample variance?3 answer-bearing and 6 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q26 · 2.7 · ProcedureHow can I fairly compare scores from two groups that have different means and spreads?5 answer-bearing and 3 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q27 · 2.7 · Numeric factWhatever the shape of the distribution, at least what share of the data lies within two standard deviations of the mean?1 answer-bearing and 2 related passages. Evidence:
Dense 9 BM25 2 Hybrid 3 |
q28 · 2.7 · Conceptual (why)Why can't I measure spread by just adding up each value's deviation from the mean?2 answer-bearing and 1 related passages. Evidence:
Dense 1 BM25 2 Hybrid 2 |
q29 · 2.7 · Conceptual (why)Why is the standard deviation less helpful for a skewed distribution, and what should I report instead?2 answer-bearing and 4 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q30 · 12.1 · DefinitionIn the textbook's form y = a + bx, which constant is the slope and which is the y-intercept?5 answer-bearing and 6 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q31 · 12.2 · DefinitionWhat is a residual, and what does a positive residual tell me?3 answer-bearing and 3 related passages. Evidence:
Dense 2 BM25 1 Hybrid 1 |
q32 · 12.2 · InterpretationHow should I interpret the slope of a least-squares line in plain words?3 answer-bearing and 3 related passages. Evidence:
Dense 5 BM25 2 Hybrid 3 |
q33 · 12.2 · InterpretationWhat does the coefficient of determination r squared tell me about a regression?3 answer-bearing and 1 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q34 · 12.2 · DefinitionWhich point does the least-squares regression line always pass through?1 answer-bearing and 0 related passages. Evidence:
Dense 16 BM25 1 Hybrid 5 |
q35 · 12.3 · ProcedureHow do I use the table of critical values to decide whether r is significant?2 answer-bearing and 9 related passages. Evidence:
Dense 2 BM25 2 Hybrid 2 |
q36 · 12.3 · Conceptual (why)What assumptions must hold for the significance test of the correlation coefficient?3 answer-bearing and 2 related passages. Evidence:
Dense 8 BM25 5 Hybrid 7 |
q37 · 12.4 · Conceptual (why)Why is it risky to predict y for an x value far outside the observed data?3 answer-bearing and 0 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q38 · 12.5 · ComparisonIn regression, how is an influential point different from an outlier?2 answer-bearing and 5 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q39 · 12.5 · ProcedureHow far from the regression line does a point have to be before I flag it as an outlier?5 answer-bearing and 3 related passages. Evidence:
Dense 1 BM25 1 Hybrid 1 |
q40 · 12.5 · Conceptual (why)Should outliers always be deleted from a regression data set?4 answer-bearing and 1 related passages. Evidence:
Dense 2 BM25 1 Hybrid 2 |
Reproduce
Recompute it, download it, check it
In your browser
Downloads about 1 MB of corpus vectors (cached afterwards) and runs every ranking, bootstrap and test again on this device.
Data and commands
uv run scripts/build_retrieval_eval.py # gold set + query vectors cd web && pnpm vitest run -u src/lib/eval # the report cd .. && uv run scripts/stats_reference.py # Python cross-check Rscript scripts/paired_ci_reference.R # R cross-check
Generation and question quality
Run the generation evaluation with your own key
Eight fixed topics through your chosen model, automatic checks on every item (citations, numbers, answerability, distractors, duplicates), your ratings on a rubric, and how often the checks agree with you. This site publishes no generation scores of its own.