Evaluation Platform
Read the transcript
1. The Math Nobody Runs Before Celebrating a Number
Host: So we’ve all seen the slide: baseline eighty-four percent, new prompt eighty-eight percent, team high-fives, ship it. I want to spend this whole episode on why that slide can be lying to you, using nothing but a fifty-example eval set. Where does the lie actually live?
Guest: It lives in the confidence interval nobody computed. Run the math on that eighty-four to eighty-eight jump on fifty examples and the ninety-five percent interval is something like negative nine point six to positive seventeen point six percent. Zero is comfortably inside that range, which means the same data is equally consistent with a ten-point regression as it is with a four-point win.
Host: So the number on the dashboard and its exact opposite are both plausible readings of the same test. What’s the actual floor, then — how big does a set need to be before a four-point move even means anything, and why does it get expensive so fast?
Guest: At an eighty-five percent baseline, fifty examples can’t resolve anything smaller than about twenty points — you’d need somewhere near two thousand examples just to reliably see a three-point delta. And the scaling is brutal because effect size enters the sample-size formula squared, so halving the delta you care about roughly quadruples the data you need — actually closer to 4.95x, since variance shrinks as the rate climbs toward one and makes the coarse measurement disproportionately cheap by comparison.
2. Teaching the Harness to Say ‘I Don’t Know’
Host: So if fifty examples usually can’t resolve anything meaningful, what does the significance function actually return in that regime? I’m guessing it doesn’t just say false.
Guest: Right, and this is the part teams miss — the significance check has three states, not two: True, False, or None. None fires when expected cell counts drop below about five, which is exactly where small eval sets live, and it means the approximation simply isn’t usable. If you collapse that into False, you’re telling someone ‘no difference’ when the honest answer is ‘this set cannot tell,’ and those two conclusions get acted on completely differently — one closes the investigation, the other should open one into whether your eval set is even fit for purpose.
Host: That’s a sharp distinction to bake into code rather than leave as a caveat in a README. But how do you know the harness itself is trustworthy enough to even return that None correctly?
Guest: That’s what the meta-test is for — you deliberately feed the harness a system that’s always wrong and assert it scores exactly 0%, plus the mirror case, a perfect system that must score 100%. If the always-wrong system doesn’t land at zero, nothing else in the suite means anything, because an eval that can’t fail isn’t an eval at all. It’s the cheapest possible check, and almost nobody runs it before trusting their scoreboard.
3. Where Eval Sets Quietly Lie, and What to Ask About Yours
Host: Okay, so beyond the meta-test, what actually goes wrong inside a real suite that a passing accuracy number won’t show you? Give me the sneaky ones.
Guest: Three come up constantly. A contains-grader will happily pass an answer that states the right fact and its opposite in the same paragraph, because the substring is technically there — verbose model output does this all the time. Flaky examples get hidden the same way: you average across repeats and get a smooth, confident-looking number instead of naming which specific examples are unstable and fixing them. And the quiet one — if your eval set was built from the system’s own past outputs, you’re measuring agreement with a former self, not correctness, and it will never catch a regression that was baked in from day one. None of that shows up in the score. You have to go looking for it.
Host: So if I’m sitting across from a team and I want to actually pressure-test their scoreboard instead of just nodding at it, what do I ask?
Guest: Four questions, in order. What’s the smallest delta this eval set can even detect — if nobody knows, nobody knows whether last quarter’s win was real. Are you pairing the tests, since running both arms on the same examples makes something like McNemar’s far more powerful, though it’s worth checking whether that pairing would’ve actually changed any decision. If you’re testing five prompt variants against one baseline, are you correcting for multiple comparisons, because that’s five chances to get lucky and it raises the sample size you actually need. And last — who verified the golden answers, because a dataset that grew out of production output is quietly measuring consistency with itself, not truth. Ask those four before you trust any number on a dashboard.
Not covered
The planner wanted these and found nothing in the source to support them:
- A worked example of McNemar’s paired test actually reducing the required sample size
- How the numbers would shift with a real (non-toy) embedding or scoring model
- Cost/latency tradeoffs of running this harness continuously vs. on a sampled cadence, per Module 4/12’s observability framing
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
An evaluation platform’s job is not to produce a number. It is to say whether a difference between two numbers is real. This lab computes what that costs in examples, and the answer is uncomfortable at the sizes teams typically use.
Source: labs/evaluation-platform
The decision path
Section titled “The decision path”flowchart TD
Base["Baseline run<br/>42 / 50 correct"] --> Cmp
Cand["Candidate run<br/>44 / 50 correct"] --> Cmp
Cmp["Comparison<br/>delta = +4.0 points"] --> Usable{"Expected cell counts >= 5?<br/>is the approximation valid here?"}
Usable -->|No| Unknown["is_significant() -> None<br/>'this set cannot tell'<br/>NOT 'no difference'"]
Usable -->|Yes| CI["95% confidence interval<br/>on the delta"]
CI --> Zero{"Does the interval contain zero?"}
Zero -->|"Yes — [-9.6%, +17.6%]"| NotReal["Not established.<br/>Same data is consistent<br/>with a 10-point regression."]
Zero -->|No| Real["Established.<br/>Report the interval, not the point."]
NotReal --> Fix["Compute what the set can see:<br/>n=50 at 85% -> 20 points minimum<br/>to see +3 points, need 2,033 per arm"]The finding
Section titled “The finding”Examples needed per arm to detect an improvement at 95% confidence and 80% power:
| Baseline | +1pt | +2pt | +3pt | +5pt | +10pt |
|---|---|---|---|---|---|
| 70% | 32,644 | 8,077 | 3,551 | 1,248 | 291 |
| 80% | 24,638 | 6,036 | 2,626 | 903 | 197 |
| 85% | 19,458 | 4,722 | 2,033 | 683 | 138 |
| 90% | 13,493 | 3,211 | 1,353 | 432 | 71 |
The inverse, which is the number worth putting on a dashboard — the smallest delta a set can see, at an 85% baseline:
| n | Smallest visible delta |
|---|---|
| 30 | 25.8% |
| 50 | 20.0% |
| 100 | 14.1% |
| 500 | 6.3% |
| 1,000 | 4.5% |
| 5,000 | 2.0% |
A fifty-example eval set cannot resolve anything below roughly twenty percentage points. Every smaller movement it reports is noise being read as signal.
The shape of most “we improved it” claims:
84% -> 88% on 50 examplesdelta +4.0%, 95% CI [-9.6%, +17.6%], significant: FalseThe same data is consistent with a ten-point regression. The point estimate on its own is not a measurement.
Effect size enters squared, so halving the delta you want to see at least quadruples the dataset — measured at 4.95× rather than the naive 4×, because as the candidate rate approaches 1.0 its variance shrinks and the coarse measurement gets disproportionately cheap.
None is a real answer
Section titled “None is a real answer”Comparison.is_significant() returns True, False, or None. The third means the
approximation is not usable — expected cell counts below about five, which is exactly the regime
small eval sets live in.
Returning False there would say no difference when the honest answer is this set cannot tell,
and those get acted on very differently: the first stops the investigation, the second should start
one about the eval set.
The meta-test
Section titled “The meta-test”test_the_harness_catches_a_deliberately_broken_system scores a system that is always wrong and
asserts it gets 0%. Everything else in the suite is worthless if that fails — an eval that cannot
fail is not an eval. Its twin asserts a perfect system scores 100%, because a harness that always
fails is equally useless.
Three things that are easy to get wrong
Section titled “Three things that are easy to get wrong”The contains grader passes an answer that says both things. An output containing the right
answer and its opposite scores as correct. That is not hypothetical — it is what verbose model
output looks like.
Flaky examples are named, not averaged away. Averaging across repeats hides instability behind a plausible number; naming the unstable examples is what lets someone fix or remove them.
An unverified dataset is visible as such. An eval set built from the system’s own past output
measures agreement with a former self, not correctness, and cannot detect a regression that was
always there. Nothing in the accuracy number reveals this, so Example carries the flag.
Run it
Section titled “Run it”cd labs/evaluation-platformuv venv .venv && uv pip install --python .venv/bin/python -e '.[dev]'./.venv/bin/python -m ruff check . && ./.venv/bin/python -m mypy src && ./.venv/bin/python -m pytest -q17 tests, ruff clean, mypy --strict clean.
Principal-level discussion points
Section titled “Principal-level discussion points”- What is your eval set’s smallest detectable delta? If nobody knows, nobody knows whether last quarter’s improvements were real.
- Paired tests would need fewer examples. Both arms run the same examples, so McNemar’s test is more powerful than the unpaired numbers here. Worth doing — and worth checking whether it changes any decision, because often it does not change the order of magnitude.
- Five prompt variants against one baseline is five chances to get lucky. Multiple-comparison correction pushes the required sample size up again.
- Who verified the golden answers? An eval set grown from production output drifts toward measuring consistency rather than correctness.
Related
Section titled “Related”- Architecture: Continuous Model Evaluation — the design review, where the sample-size tables above become the platform’s largest line item.
- Module 4: AI Infrastructure — where evaluation sits in the cost-latency-quality triangle.
- Module 12: Observability — measuring a system you cannot see inside.
- Semantic Cache — an eval suite running through a cache is measuring the cache, not the model.
hybrid-retrieval— retrieval-specific evaluation, where retrieval quality and answer quality are measured separately.