Skip to content
Lecture Revision Lab

Evaluation

Does it find the right passage, and can you trust what the model writes?

Retrieval is measured on 40 questions about the demo textbook, comparing the 2024 retriever with BM25 and a hybrid on the same questions. Every number has a 95% interval, every comparison is paired, and the whole report can be recomputed in your browser. Generation is evaluated with your own key.

Headline results

Provisional: the questions and relevance labels were drafted by Claude Opus 5.5 under the owner's direction and have not yet been reviewed by a person. The numbers will be recomputed after that review.

Highest nDCG@10 (point estimate)provisional

Hybrid 0.794

95% interval 0.745 to 0.839. Not distinguishable from BM25 (0.778; difference +0.016, −0.040 to 0.066); dense 0.746.

Hybrid minus the 2024 retrieverprovisional

+0.048

Paired interval 0.008 to 0.089; Holm-adjusted p = 0.084, so not significant by the rule declared before the first run.

BM25 answer in the top 5provisional

40 of 40

Wilson interval 0.91 to 1.00; dense 36 of 40, hybrid 39 of 40.

Chunks over whole sectionsprovisional

+0.25

Answering section ranked first, 36 vs 26 of 40; paired score interval 0.07 to 0.42; McNemar exact p = 0.013.

How it was measured

A pooled gold set, judged before any score was computed

Full design in DR-003; assumptions and limitations on Methods.
Questions
40, in six types, written as a student would ask; the app's sample queries are excluded.
Relevance
119 answer-bearing passages (hold a verbatim evidence span) and 135 related ones.
Pooling
Top 10 of all three rankers, 480 judgements drafted by Claude Opus 5.5 blind to the ranker, not yet reviewed by a person. Unjudged passages count as not relevant. For the chunking comparison, every section in either arm's top 3 was read too (128 question-section pairs).
Statistics
Bootstrap 10,000 resamples, seed 20241009; Wilson for hit rates; sign-flip test (20,000 sign vectors when not exact); Holm over the three primary comparisons.
  • MiniLM cosine (the 2024 retriever)all-MiniLM-L6-v2 vectors and a dot product of unit vectors: the prototype's CustomRM scoring, unchanged.
  • BM25 keyword searchOkapi BM25 (k1 = 1.2, b = 0.75) over lower-cased, stop-word-filtered, lightly stemmed tokens: the app's offline fallback.
  • Hybrid (reciprocal rank fusion)Reciprocal rank fusion of the dense and BM25 top 100 with k = 60, the value from the original paper.

Results by ranker

nDCG@10 point estimates: Hybrid 0.794, BM25 0.778, Dense 0.746

Only hybrid minus dense has an interval excluding zero, and it does not survive Holm's adjustment. Dots are estimates over the 40 questions; whiskers are 95% intervals (bootstrap for means, Wilson for hit rates). nDCG@10 is the primary metric. The axes start at 0.5.
  • Dense (circle)
  • BM25 (square)
  • Hybrid (diamond)
nDCG@10 (primary)
Dense0.746 (0.674 to 0.811)
BM250.778 (0.722 to 0.829)
Hybrid0.794 (0.745 to 0.839)
MRR@10
Dense0.750 (0.642 to 0.854)
BM250.827 (0.740 to 0.908)
Hybrid0.761 (0.664 to 0.853)
Recall@5
Dense0.677 (0.567 to 0.784)
BM250.729 (0.655 to 0.802)
Hybrid0.755 (0.671 to 0.833)
Success@5 (answer in the top 5)
Dense0.900 (0.769 to 0.960)
BM251.000 (0.912 to 1.000)
Hybrid0.975 (0.871 to 0.996)
All retrieval metrics by ranker with 95% intervals
MetricDenseBM25Hybrid
nDCG@10primary0.746 (0.674 to 0.811)0.778 (0.722 to 0.829)0.794 (0.745 to 0.839)
MRR@100.750 (0.642 to 0.854)0.827 (0.740 to 0.908)0.761 (0.664 to 0.853)
Recall@50.677 (0.567 to 0.784)0.729 (0.655 to 0.802)0.755 (0.671 to 0.833)
Recall@100.820 (0.722 to 0.906)0.856 (0.789 to 0.918)0.877 (0.810 to 0.935)
Success@125 / 40 (0.47 to 0.76)28 / 40 (0.55 to 0.82)24 / 40 (0.45 to 0.74)
Success@334 / 40 (0.71 to 0.93)38 / 40 (0.83 to 0.99)37 / 40 (0.80 to 0.97)
Success@536 / 40 (0.77 to 0.96)40 / 40 (0.91 to 1.00)39 / 40 (0.87 to 1.00)
Success@1038 / 40 (0.83 to 0.99)40 / 40 (0.91 to 1.00)40 / 40 (0.91 to 1.00)

Paired comparisons

Hybrid minus dense: +0.048, not significant after Holm's adjustment

Each difference is computed on the same questions and resampled in pairs. On the primary metric the hybrid beats dense on 28 questions and loses on 11 (d_z = 0.36, a small effect). Its interval excludes zero, and after Holm's adjustment for the three planned comparisons p = 0.084, so it is not a finding. With 40 questions, a difference of a few hundredths cannot be told from noise.
Paired difference in nDCG@10 (A minus B)
Hybrid − Dense+0.048 (0.008 to 0.089)
Hybrid − BM25+0.016 (−0.040 to 0.066)
Dense − BM25−0.032 (−0.121 to 0.050)
Paired comparisons with intervals, effect sizes and tests
MetricA − BDifference (95% interval)d_zWins / ties / lossesp (sign-flip)Holm pMcNemar p
nDCG@10primaryHybrid − Dense+0.048 (0.008 to 0.089)bootstrap0.3628 / 1 / 110.028MC0.084–
nDCG@10primaryHybrid − BM25+0.016 (−0.040 to 0.066)bootstrap0.0920 / 1 / 190.559MC0.968–
nDCG@10primaryDense − BM25−0.032 (−0.121 to 0.050)bootstrap−0.1115 / 0 / 250.484MC0.968–
MRR@10Hybrid − Dense+0.010 (−0.060 to 0.080)bootstrap0.058 / 28 / 40.792exact––
MRR@10Hybrid − BM25−0.066 (−0.157 to 0.021)bootstrap−0.234 / 26 / 100.162exact––
MRR@10Dense − BM25−0.076 (−0.199 to 0.041)bootstrap−0.197 / 21 / 120.244MC––
Success@5Hybrid − Dense+0.075 (−0.033 to 0.203)score0.283 / 37 / 00.250exact–0.250
Success@5Hybrid − BM25−0.025 (−0.129 to 0.065)score−0.160 / 39 / 11.000exact–1.000
Success@5Dense − BM25−0.100 (−0.231 to 0.006)score−0.330 / 36 / 40.125exact–0.125

Secondary rows are descriptive: their p-values are not adjusted and do not count as findings. Ties are questions where both rankers scored the same, which is common for hit-based metrics. Differences in means use the paired bootstrap; Success@5 differences use Newcombe's paired score interval, because with a handful of discordant questions all pointing one way the bootstrap interval is too narrow (for dense − BM25 it was −0.200 to −0.025, excluding zero, while McNemar's exact p was 0.125).

Indexing unit

Chunks find the answering section far more often than whole sections

Same model, same cosine scoring; only the unit changes. The 2024 index has one vector per section, made from its first 256 word pieces. Chunks are collapsed to their sections for a fair comparison, and every section either arm ranked in its top 3 was read for every question, so no miss comes from a section nobody judged (128 question-section pairs, read after review; none held a further answer). The difference in section-level Success@1 is +0.250 (paired score interval 0.068 to 0.415), McNemar exact p = 0.013 (12 questions only chunks got right, 2 only whole sections did). Reasoning in DR-002.
Answering section in the top 1
Whole sections (2024)26 of 40 (0.50 to 0.78)
Chunks (2026)36 of 40 (0.77 to 0.96)
Answering section in the top 3
Whole sections (2024)35 of 40 (0.74 to 0.95)
Chunks (2026)40 of 40 (0.91 to 1.00)

Exploratory

By question type

Three to thirteen questions per type, so these intervals are wide and the breakdown is for curiosity, not conclusions.
nDCG@10 and Success@5 by question type
TypenDense nDCG@10 · S@5BM25 nDCG@10 · S@5Hybrid nDCG@10 · S@5
Definition50.59 (0.24 to 0.91) · 3/50.86 (0.76 to 0.95) · 5/50.75 (0.56 to 0.91) · 5/5
Procedure90.83 (0.72 to 0.92) · 9/90.80 (0.74 to 0.87) · 9/90.86 (0.82 to 0.90) · 9/9
Interpretation60.77 (0.57 to 0.92) · 6/60.70 (0.54 to 0.86) · 6/60.80 (0.68 to 0.90) · 6/6
Comparison40.73 (0.57 to 0.88) · 4/40.83 (0.73 to 0.93) · 4/40.82 (0.71 to 0.94) · 4/4
Conceptual (why)130.73 (0.64 to 0.81) · 12/130.76 (0.65 to 0.86) · 13/130.75 (0.66 to 0.84) · 12/13
Numeric fact30.81 (0.71 to 0.98) · 2/30.75 (0.40 to 0.95) · 3/30.81 (0.58 to 1.00) · 3/3

Every question

Per-question rankings

Each strip shows a ranker's top 10: a filled square is an answer-bearing passage, a half-filled one is related, an empty one was judged not relevant. The number is the rank of the first answer-bearing passage (in the top 100). Open a question to see its evidence.
Top 10 of each ranker for every gold question
Question and each ranker's top 10
q01 · 1.1 · ComparisonI surveyed 100 students to estimate the average for every student at the school. Which number is the parameter and which is the statistic?

6 answer-bearing and 4 related passages. Evidence:

  • “A statistic is a number that represents a property of the sample.”
  • “A parameter is a numerical characteristic of the whole population that can be estimated by a statistic.”
  • “The parameter is the mean amount of extracurricular activities in which all high school students participate.”
  • “Key term. statistic: a numerical characteristic of the sample; a statistic estimates the corresponding population parameter”
  • “The sample mean is a statistic.”
  • “The population mean is a parameter.”
Dense
BM25
Hybrid
q02 · 1.1 · Conceptual (why)Does it make sense to calculate an average for a categorical variable such as party affiliation?

4 answer-bearing and 3 related passages. Evidence:

  • “calculating an average party affiliation makes no sense”
  • “For example, it does not make sense to find an average hair color or blood type.”
  • “Nominal scale data cannot be used in calculations.”
Dense
BM25
Hybrid
q03 · 1.2 · ProcedureHow can I tell whether quantitative data are discrete or continuous?

4 answer-bearing and 6 related passages. Evidence:

  • “All data that are the result of counting are called quantitative discrete data.”
  • “Data that are not only made up of counting numbers, but that may include fractions, decimals, or irrational numbers, are called quantitative continuous data.”
  • “Data is discrete if it is the result of counting”
  • “are quantitative continuous data because you measure weights as precisely as possible.”
Dense
BM25
Hybrid
q04 · 1.2 · ComparisonWhen should I use a bar graph rather than a pie chart for categorical data?

4 answer-bearing and 5 related passages. Evidence:

  • “We use pie charts when we want to show parts of a whole.”
  • “A bar graph is appropriate to compare the relative size of the categories. A pie chart cannot be used.”
  • “In this situation, create a bar graph and not a pie chart.”
  • “They should use the bar graph if knowing the exact numbers of students and the relative sizes of each category at each school are important points to be made.”
Dense
BM25
Hybrid
q05 · 1.2 · DefinitionIf I randomly pick a few whole classrooms and survey every student in them, what kind of sample is that?

3 answer-bearing and 2 related passages. Evidence:

  • “To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters.”
  • “the four classes make up the cluster sample”
  • “Key term. cluster sampling: a method for selecting a random sample and dividing the population into groups (clusters)”
Dense
BM25
Hybrid
q06 · 1.2 · ComparisonWhat is the difference between sampling error and nonsampling error?

3 answer-bearing and 3 related passages. Evidence:

  • “The actual process of sampling causes sampling errors.”
  • “Factors not related to the sampling process cause nonsampling errors.”
  • “Key term. nonsampling error: an issue that affects the reliability of sampling data other than natural variation”
  • “Key term. sampling error: the natural variation that results from selecting a sample to represent a larger population”
Dense
BM25
Hybrid
q07 · 1.2 · Conceptual (why)Why are results from online polls that anyone can answer often unreliable?

4 answer-bearing and 2 related passages. Evidence:

  • “Self-selected samples—Responses only by people who choose to respond, such as internet surveys, are often unreliable.”
  • “internet surveys are invariably biased, because people choose to respond or not.”
  • “Surveys mailed to households and then returned may be very biased.”
Dense
BM25
Hybrid
q08 · 1.3 · Conceptual (why)Why is temperature in degrees Celsius measured on an interval scale and not a ratio scale?

2 answer-bearing and 3 related passages. Evidence:

  • “Temperature scales like Celsius (C) and Fahrenheit (F) are measured by using the interval scale.”
  • “But 0 degrees does not represent a minimum value.”
  • “Interval scale level data with a definite ordering but no starting point; the differences can be measured, but there is no such thing as a ratio”
Dense
BM25
Hybrid
q09 · 1.3 · ProcedureHow do I fill in the cumulative relative frequency column of a frequency table?

4 answer-bearing and 5 related passages. Evidence:

  • “To find the cumulative relative frequencies, add all the previous relative frequencies to the relative frequency for the current row”
  • “To find the cumulative relative frequency, add all of the previous relative frequencies to the relative frequency for the current row.”
  • “The cumulative relative frequency is the sum of the relative frequencies for all values that are less than or equal to the given value”
  • “Continue adding the relative frequencies in each row to get the rest of the column.”
Dense
BM25
Hybrid
q10 · 1.4 · Conceptual (why)Why does randomly assigning subjects to treatment groups let an experiment show cause and effect?

6 answer-bearing and 2 related passages. Evidence:

  • “When subjects are assigned treatments randomly, all of the potential lurking variables are spread equally among the groups.”
  • “In this way, an experiment can prove a cause-and-effect connection between the explanatory and response variables.”
  • “Random assignment eliminates the problem of lurking variables.”
  • “Random assignment eliminates the impact of lurking variables.”
  • “To eliminate lurking variables, subjects must be assigned randomly to different treatment groups.”
  • “When a study is designed properly, the only difference between treatment groups is the one imposed by the researcher.”
Dense
BM25
Hybrid
q11 · 1.4 · Conceptual (why)Why do experiments give the control group a placebo?

3 answer-bearing and 5 related passages. Evidence:

  • “To counter the power of suggestion, researchers set aside one treatment group as a control group.”
  • “This group is given a placebo treatment, a treatment that cannot influence the response variable.”
  • “Participants in the control group receive a placebo treatment that looks exactly like the active treatments but cannot influence the response variable.”
Dense
BM25
Hybrid
q12 · 1.4 · Conceptual (why)When would a researcher run an observational study instead of an experiment?

1 answer-bearing and 3 related passages. Evidence:

  • “Sometimes, it is neither possible nor ethical for researchers to conduct experimental studies.”
  • “In these cases, observational studies or surveys may be used.”
Dense
BM25
Hybrid
q13 · 2.1 · Numeric factIn a stem-and-leaf plot, what are the stem and the leaf of the number 432?

1 answer-bearing and 2 related passages. Evidence:

  • “The number 432 has stem 43 and leaf two.”
  • “The stem consists of the leading digit(s), while the leaf consists of a final significant digit.”
Dense
BM25
Hybrid
q14 · 2.2 · ProcedureHow do I decide how many bars a histogram should have and how wide each bar should be?

3 answer-bearing and 2 related passages. Evidence:

  • “Many histograms consist of five to 15 bars or classes for clarity.”
  • “which may be calculated by dividing the range of the data values by the desired number of bins (or bars).”
  • “A guideline that is followed by some for the width of a bar or class interval is to take the square root of the number of data values and then round to the nearest whole number, if necessary.”
Dense
BM25
Hybrid
q15 · 2.2 · ProcedureIf a value lands exactly on the boundary between two histogram classes, which class does it belong to?

1 answer-bearing and 3 related passages. Evidence:

  • “A value is counted in a class interval if it falls on the left boundary but not if it falls on the right boundary.”
Dense
BM25
Hybrid
q16 · 2.3 · InterpretationIf I scored in the 90th percentile, did I get 90 percent of the questions right?

2 answer-bearing and 6 related passages. Evidence:

  • “To score in the 90^th percentile of an exam does not mean, necessarily, that you received 90 percent on a test.”
  • “It means that 90 percent of test scores are the same as or less than your score”
Dense
BM25
Hybrid
q17 · 2.3 · ProcedureHow far from the quartiles does a value have to be before it counts as a potential outlier?

2 answer-bearing and 4 related passages. Evidence:

  • “A value is suspected to be a potential outlier if it is less than 1.5 × IQR below the first quartile or more than 1.5 × IQR above the third quartile.”
  • “The IQR is found by subtracting Q_1 from Q_3 and can help determine outliers by using the following two expressions.”
Dense
BM25
Hybrid
q18 · 2.4 · DefinitionWhat do the whiskers of a box plot show?

3 answer-bearing and 3 related passages. Evidence:

  • “The whiskers extend from the ends of the box to the smallest and largest data values.”
  • “The two whiskers extend from the first quartile to the smallest value and from the third quartile to the largest value.”
Dense
BM25
Hybrid
q19 · 2.4 · InterpretationWhen comparing two box plots, what does it mean if one box is much wider than the other?

1 answer-bearing and 5 related passages. Evidence:

  • “The IQR for the first data set is greater than the IQR for the second set. This means that there is more variability in the middle 50 percent of the first data set.”
Dense
BM25
Hybrid
q20 · 2.5 · InterpretationIn a town where one person earns millions and everyone else earns about the same, should I report the mean or the median income?

3 answer-bearing and 3 related passages. Evidence:

  • “The median is generally a better measure of the center when there are extreme values or outliers because it is not affected by the precise numerical values of the outliers.”
  • “The median is a better measure of the center than the mean because 49 of the values are 30,000 and one is 5,000,000.”
  • “the median is the best measurement when a data set contains several outliers or extreme values”
Dense
BM25
Hybrid
q21 · 2.5 · ProcedureHow can I estimate the mean from a grouped frequency table when the individual values are unknown?

5 answer-bearing and 3 related passages. Evidence:

  • “Since we do not know the individual data values, we can instead find the midpoint of each interval.”
  • “So this formula says that we will sum the products of each midpoint and the corresponding frequency and divide by the sum of all of the frequencies.”
  • “the mean can be approximated if you add the lower boundary with the upper boundary and divide by two to find the midpoint of each interval.”
  • “determine the best estimate of the measures of center by finding the mean of the grouped data with the formula”
Dense
BM25
Hybrid
q22 · 2.5 · Numeric factWith 100 data values, where in the ordered list is the median?

2 answer-bearing and 2 related passages. Evidence:

  • “The median occurs midway between the 50^th and 51^st values.”
  • “You can quickly find the location of the median by using the expression (n + 1)/2 .”
Dense
BM25
Hybrid
q23 · 2.5 · Conceptual (why)What happens to the sample mean as I take bigger and bigger samples?

2 answer-bearing and 6 related passages. Evidence:

  • “The Law of Large Numbers says that if you take samples of larger and larger size from any population”
  • “their sample results (the average amount of time a student sleeps) might be closer to the actual population average.”
Dense
BM25
Hybrid
q24 · 2.6 · InterpretationIf the mean of my data is bigger than the median, which way is the distribution skewed?

3 answer-bearing and 4 related passages. Evidence:

  • “If the distribution of data is skewed to the right, the mode is often less than the median, which is less than the mean.”
  • “The mean is pulled toward the tail in a skewed distribution.”
  • “In a skewed distribution, the mean is pulled toward the tail of the distribution.”
Dense
BM25
Hybrid
q25 · 2.7 · Conceptual (why)Why do we divide by n minus 1 when calculating a sample variance?

3 answer-bearing and 6 related passages. Evidence:

  • “dividing by (n – 1) gives a better estimate of the population variance.”
  • “If the data are from a sample rather than a population, when we calculate the average of the squared deviations, we divide by n – 1, one less than the number of items in the sample.”
Dense
BM25
Hybrid
q26 · 2.7 · ProcedureHow can I fairly compare scores from two groups that have different means and spreads?

5 answer-bearing and 3 related passages. Evidence:

  • “a z-score allows us to compare statistics from different data sets.”
  • “If the data sets have different means and standard deviations, then comparing the data values directly can be misleading.”
  • “A z-score is a standardized score that lets us compare data sets.”
  • “For each student, determine how many standard deviations (#ofSTDEVs) his GPA is away from the average, for his school.”
Dense
BM25
Hybrid
q27 · 2.7 · Numeric factWhatever the shape of the distribution, at least what share of the data lies within two standard deviations of the mean?

1 answer-bearing and 2 related passages. Evidence:

  • “At least 75 percent of the data is within two standard deviations of the mean.”
  • “This is known as Chebyshev's Rule.”
Dense
BM25
Hybrid
q28 · 2.7 · Conceptual (why)Why can't I measure spread by just adding up each value's deviation from the mean?

2 answer-bearing and 1 related passages. Evidence:

  • “If you add the deviations, the sum is always zero.”
  • “So you cannot simply add the deviations to get the spread of the data.”
  • “By squaring the deviations, you make them positive numbers, and the sum will also be positive.”
Dense
BM25
Hybrid
q29 · 2.7 · Conceptual (why)Why is the standard deviation less helpful for a skewed distribution, and what should I report instead?

2 answer-bearing and 4 related passages. Evidence:

  • “in skewed distributions, the standard deviation may not be much help.”
  • “In a skewed distribution, it is better to look at the first quartile, the median, the third quartile, the smallest value, and the largest value.”
Dense
BM25
Hybrid
q30 · 12.1 · DefinitionIn the textbook's form y = a + bx, which constant is the slope and which is the y-intercept?

5 answer-bearing and 6 related passages. Evidence:

  • “b = slope and a = y-inttercept”
  • “In this text, the form y = a + bx is used, where a is the y-intercept and b is the slope.”
  • “In the equation y = a + bx, the constant b that multiplies the x variable (b is called a coefficient) is called the slope.”
  • “The y-intercept is 25 (a = 25).”
  • “In the equation y = a + bx, the constant a is called the y-intercept.”
Dense
BM25
Hybrid
q31 · 12.2 · DefinitionWhat is a residual, and what does a positive residual tell me?

3 answer-bearing and 3 related passages. Evidence:

  • “is called the error or residual”
  • “If the observed data point lies above the line, the residual is positive and the line underestimates the actual data value for y.”
  • “Residuals, also called errors, measure the distance from the actual value of y and the estimated value of y.”
Dense
BM25
Hybrid
q32 · 12.2 · InterpretationHow should I interpret the slope of a least-squares line in plain words?

3 answer-bearing and 3 related passages. Evidence:

  • “The slope of the best-fit line tells us how the dependent variable (y) changes for every one unit increase in the independent (x) variable, on average.”
  • “The slope tells us how the dependent variable (y) changes for every one-unit increase in the independent (x) variable, on average.”
  • “For a 1-point increase in the score on the third exam, the final exam score increases by 4.83 points, on average.”
  • “the rate of change describes the change that occurs in the dependent variable as the independent variable is changed.”
Dense
BM25
Hybrid
q33 · 12.2 · InterpretationWhat does the coefficient of determination r squared tell me about a regression?

3 answer-bearing and 1 related passages. Evidence:

  • “represents the percentage of variation in the dependent (predicted) variable y that can be explained by variation in the independent (explanatory) variable x using the regression (best-fit) line.”
  • “r^2 represents the percentage of variation in the dependent variable, y, that can be explained by variation in the independent variable, x, using the regression line.”
  • “Approximately 44 percent of the variation (0.4397 is approximately 0.44) in the final exam grades can be explained by the variation in the grades on the third exam”
Dense
BM25
Hybrid
q34 · 12.2 · DefinitionWhich point does the least-squares regression line always pass through?

1 answer-bearing and 0 related passages. Evidence:

  • “The best-fit line always passes through the point”
Dense
BM25
Hybrid
q35 · 12.3 · ProcedureHow do I use the table of critical values to decide whether r is significant?

2 answer-bearing and 9 related passages. Evidence:

  • “Use it to find the critical values using the degrees of freedom, df = n – 2.”
  • “If r is not between the positive and negative critical values, then the correlation coefficient is significant.”
Dense
BM25
Hybrid
q36 · 12.3 · Conceptual (why)What assumptions must hold for the significance test of the correlation coefficient?

3 answer-bearing and 2 related passages. Evidence:

  • “The y values for any particular x value are normally distributed about the line.”
  • “The residual errors are mutually independent (no pattern).”
  • “Equal variance: The standard deviation of the y values is equal for each x value.”
Dense
BM25
Hybrid
q37 · 12.4 · Conceptual (why)Why is it risky to predict y for an x value far outside the observed data?

3 answer-bearing and 0 related passages. Evidence:

  • “outside the domain of the observed x values in the data (independent variable), so you cannot reliably predict”
  • “should not be used to make predictions for values outside the set of data.”
  • “the line may not be appropriate or reliable for prediction outside the domain of observed x values in the data.”
Dense
BM25
Hybrid
q38 · 12.5 · ComparisonIn regression, how is an influential point different from an outlier?

2 answer-bearing and 5 related passages. Evidence:

  • “Influential points are observed data points that are far from the other observed data points in the horizontal direction.”
  • “Outliers are observed data points that are far from the least-squares line.”
  • “These points may have a big effect on the slope of the regression line.”
Dense
BM25
Hybrid
q39 · 12.5 · ProcedureHow far from the regression line does a point have to be before I flag it as an outlier?

5 answer-bearing and 3 related passages. Evidence:

  • “we can flag as an outlier any point that is located farther than two standard deviations above or below the best-fit line.”
  • “If the absolute value of any residual is greater than or equal to 2s, then the corresponding point is an outlier.”
  • “that distance were equal to 2s or more, then we would consider the data point to be too far from the line of best fit.”
  • “it is considered an outlier because it is more than two standard deviations away from the best-fit line.”
  • “that distance is at least 2s, then we would consider the data point to be too far from the line of best fit.”
Dense
BM25
Hybrid
q40 · 12.5 · Conceptual (why)Should outliers always be deleted from a regression data set?

4 answer-bearing and 1 related passages. Evidence:

  • “Other times, an outlier may hold valuable information about the population under study and should remain included in the data.”
  • “(Remember, we do not always delete an outlier.)”
  • “When outliers are deleted, the researcher should either record that data were deleted, and why, or the researcher should provide results both with and without the deleted data.”
  • “If the outlier was recorded erroneously, it should certainly be deleted.”
Dense
BM25
Hybrid

Reproduce

Recompute it, download it, check it

The report is built by deterministic code from the corpus, the gold set and stored query vectors. A test keeps it identical to the code, a Python script recomputes every metric, interval and test (dense and section rankings from the stored vectors) with numpy, scipy, statsmodels and pytrec_eval, and an R script checks the paired score intervals with contingencytables.

In your browser

Downloads about 1 MB of corpus vectors (cached afterwards) and runs every ranking, bootstrap and test again on this device.

Data and commands

uv run scripts/build_retrieval_eval.py   # gold set + query vectors
cd web && pnpm vitest run -u src/lib/eval  # the report
cd .. && uv run scripts/stats_reference.py # Python cross-check
Rscript scripts/paired_ci_reference.R      # R cross-check

Generation and question quality

Run the generation evaluation with your own key

Eight fixed topics through your chosen model, automatic checks on every item (citations, numbers, answerability, distractors, duplicates), your ratings on a rubric, and how often the checks agree with you. This site publishes no generation scores of its own.

Open the harness