Architecture: Continuous Model Evaluation
Read the transcript
1. The number that isn’t a decision
Host: So let’s start with the situation basically every AI team is living in right now. Someone tweaks a prompt, swaps a model, changes a retrieval strategy, and they run it through an eval harness, and out pops something like eighty-four percent going to eighty-eight percent. And the instinct is to treat that as good news.
Guest: Right, and that instinct is exactly the trap. Eighty-four to eighty-eight isn’t a decision, it’s just two numbers sitting next to each other. Whether that four-point jump means the change actually helped, or it’s just noise from which examples happened to be in the set, or it’s even hiding a real regression underneath — you literally cannot tell from the numbers alone, especially at the dataset sizes most teams are actually running.
Host: So the platform did its job, technically — it produced a score — but it never answered the question anyone actually cared about, which is ‘should I ship this.’ What has to be true for that number to actually mean something?
Guest: Three things, and each one can fail without anyone noticing. The grader has to be capable of marking something wrong, otherwise a broken system and a working one look identical. The dataset has to be big enough to even see the size of improvement you’re claiming, since small deltas are invisible below a certain scale. And the dataset has to be checked against actual correctness, not just against what the system itself produced last time — otherwise you’re just measuring agreement with your own past mistakes.
2. Three conditions that fail silently
Host: So walk me through why these three things fail silently. If a grader can’t fail, what does that actually look like in practice — is it obviously broken, or does it just look fine on the dashboard?
Guest: It looks completely fine, that’s the problem. You’ll see a harness that scores everything highly because the grading prompt is too lenient, or the rubric only checks for surface features like ‘did it mention the right keyword.’ Nobody gets an error message. You just get a green number every single time, for a broken system and a good one alike, and the dashboard never tells you which one you’re looking at.
Host: And the same goes for dataset size and the correctness check — there’s no alarm bell, just a number that quietly stops meaning anything.
Guest: Exactly. At the dataset sizes teams actually use, a set will happily report a four-point swing and nobody can tell whether that swing is real signal or just noise. And if your ‘ground truth’ is just last month’s model output, you’re not measuring correctness at all — you’re measuring how well the new system agrees with the old one’s mistakes, which rewards consistency, not accuracy. All three failures produce a perfectly normal-looking number, which is exactly why they get missed.
3. Turning the conditions into a checklist
Host: So let’s turn those three failure modes into something you can actually check for in a platform. If I hand you an eval system, what are the concrete boxes it needs to tick?
Guest: First, the grader itself has to pass a meta-test: feed it a deliberately broken system and it must score zero, feed it a perfect one and it must score a hundred. Then every comparison reports an interval, never a bare point estimate, and each set publishes its own smallest detectable delta so a claim can be checked against what that set can actually resolve. On top of that you want provenance per example — who verified this answer and against what — flaky examples named instead of averaged away, and a path to the model that skips the response cache entirely.
Host: That cache one is easy to miss — walk me through why bypassing it matters as much as the statistics do.
4. Why effect size squared breaks everyone’s budget
Host: We’ll get to the cache, but first I want to sit with a number you keep coming back to: effect size enters the sample-size calculation squared. Why does that one detail wreck so many teams’ evaluation budgets?
Guest: Because it means intuition about cost is wrong by an order of magnitude. Halving the delta you want to detect more than quadruples the examples required — effect size enters the calculation squared. At an 85% baseline, going from a 5-point detectable delta to a 3-point one takes you from 683 examples to over 2,000. People plan evaluation sets like the relationship is linear, and it just isn’t.
Host: So walk me through what that actually costs someone. If a team wants to catch a three-point improvement at an 85% baseline, what are they signing up for?
Guest: Roughly two thousand verified examples per arm — and that’s the architecture’s biggest line item, because model spend is examples times arms times repeats, and curation time scales the same way. Flip it around and the smallest-visible-delta table tells the uncomfortable truth: a 50-example set can only see a 20-point swing, so reporting ‘88%’ without also reporting that resolution is how a four-point improvement gets treated as real when the set was structurally blind to it.
5. Anatomy of a false improvement: 84% to 88%
Host: So walk me through the canonical case, the one everybody’s seen without realizing it: 84% goes to 88% on a 50-example set. On its face that’s a win — where does it actually fall apart?
Guest: It falls apart the moment you compute the interval instead of just the point estimate. That four-point gain carries a 95% confidence interval running from about negative 9.6% to positive 17.6% — which means the exact same data is fully consistent with a ten-point regression. You didn’t measure an improvement, you measured a number that improvement is compatible with, along with about a dozen other stories including ‘this got worse.’
Host: And nobody reports the interval, they just report the 88 and move on.
Guest: Right, and that’s how it becomes roadmap-worthy — someone screenshots ‘84 to 88,’ it goes in a slide, and a quarter of engineering effort gets justified by four points that a 50-example set was never built to see in the first place. The set’s own resolution floor at that baseline is roughly twenty points, so this measurement was noise dressed up as signal from the moment it was collected.
6. Undecidable is a verdict, not a shrug
Host: So that resolution-floor problem you just described — is that what forces the three-state design? Most people would just say the comparison returns true or false, significant or not.
Guest: Exactly that. Comparison dot is_significant returns True, False, or None, and None isn’t a bug, it’s the honest answer when expected cell counts drop below about five — which is precisely the regime small eval sets live in. If you force that into False, you’re not simplifying, you’re lying, because False says ‘no difference’ and None says ‘this set cannot tell,’ and those two claims should never be acted on the same way.
Host: So collapsing them is the actual failure mode — not a rounding error, but a category error that kills the investigation before it starts.
Guest: Right, because False closes the case — engineers move on, nothing to see here. None should open a case, but about the eval set itself: why can’t this instrument resolve a question this important, and what would it take to fix that. Undecidable is a verdict that points the investigation somewhere; false is a verdict that ends it.
7. The meta-test: can your eval even fail?
Host: So before we trust any of these numbers — the effect sizes, the confidence intervals, all of it — is there a test underneath all of them, something that checks the checker itself?
Guest: There’s exactly one test that has to exist before any other number means anything: feed the harness a system that’s always wrong, and assert it scores zero. If that fails, every other result in the suite is unverified, because you have no proof the grader can distinguish wrong from right — it might just be returning correct on an exception and calling it a pass. The twin test asserts a perfect system scores 100%, because a harness that always fails is just as useless as one that never does; an eval that cannot fail, in either direction, is not an eval.
8. Graders are systems too — and need their own evaluation
Host: So the harness itself can pass that zero-and-hundred test and still lie to you, because the grader sitting inside it is broken in a different way. Walk me through the contains grader problem — you said it passes an output that says both things?
Guest: Right, a contains grader just checks whether the right answer appears as a substring, and verbose model output routinely states the correct answer and its opposite in the same response — hedging, showing both sides, restating the wrong premise before correcting it. The grader sees the right string, doesn’t see that it’s sitting next to its negation, and marks it correct. That’s not an edge case, that’s what a chatty model does by default, so a contains grader can pass exactly the kind of output that a careful reader would call wrong.
Host: And that’s presumably why people reach for an LLM-as-judge instead — let a model read the whole thing and decide. But you’re telling me that’s the least trustworthy piece in the entire pipeline?
Guest: It’s the least verifiable component, full stop, because it needs its own evaluation against its own verified set before its verdicts mean anything — you can’t just trust a judge model because it sounds reasonable. And there’s a sharper problem underneath that: the judge is a model reading text that the system under test produced, which means it’s an injection surface. If a graded output can address the judge directly, the system under test can influence its own score, and now you don’t have an evaluation, you have a conversation the candidate is winning.
9. When the dataset measures agreement with itself
Host: So beyond graders lying, there’s a problem with the dataset itself. You mentioned earlier that some eval sets grow from production output — walk me through how that goes wrong.
Guest: Say you don’t have a human-verified answer for some tricky input, so you take what the system produced last quarter and label that the expected answer. Fine, once. But do that repeatedly and the set stops measuring correctness — it starts measuring whether the system still agrees with its earlier self. If the system was wrong back then, it’ll be wrong forever and score perfectly, because the wrong answer is now the target it’s being compared against.
Host: And the accuracy number gives no hint that’s happening.
Guest: None. It looks entirely normal. That’s why provenance can’t be an assumption you carry in your head about how the set was built — it has to be a recorded field on every example, human-verified versus captured-from-output, so you can report what fraction of your ‘accuracy’ is actually measuring agreement with a possibly-broken past rather than correctness.
10. Naming instability instead of averaging it away
Host: So say you run each example five times to smooth out noise, and average the results. That sounds responsible — where does it go wrong?
Guest: It goes wrong because averaging makes two very different problems look identical. An example that passes half the time and one that deterministically gets half credit both show up as 50% — but one is a coin flip you need to fix or throw out, and the other is a stable partial-match you can reason about. The fix is to name the flaky ones individually — flag which examples are unstable across repeats — so someone can actually look at them, instead of burying that signal inside a plausible-looking average that never tells you which examples are the problem.
11. The cache hiding underneath the eval
Host: You mentioned the cache in passing a minute ago, and I want to stop on it, because it sounds like the eval suite can fail in a way that has nothing to do with any of the statistical problems we’ve covered. What’s going on there?
Guest: Right, this is a completely different failure class. If your eval suite calls the model through the same path production traffic uses, and that path has a semantic cache sitting in front of the model, your suite isn’t measuring the model — it’s measuring whatever the cache decided to return. So you ship a model upgrade, run your eval, and the score is identical to last time. That reads as ‘the upgrade did nothing,’ but what actually happened is a bunch of your eval examples hit cached responses from the old model and never reached the new one at all.
Host: So the eval passes, the dashboard looks stable, and there’s just no signal in it at all.
Guest: None — and it’s worse than a neutral non-signal, because it actively looks like confirmation that nothing broke. That’s why bypass can’t be an option you remember to set; it has to be the default for anything calling itself an evaluation. Every other fix we’ve talked about — the intervals, the meta-test, naming instability — assumes the number you’re looking at came from the system you think it did, and the cache is the one place that assumption can be false without a single line of your eval code being wrong.
12. What it costs, who can touch it, and what to watch
Host: Let’s put a number on all of this, because I think people hear ‘run more examples’ and don’t feel the actual cost. If you want to detect a three-point improvement at an 85% baseline, what are you signing up for?
Guest: Roughly two thousand verified examples per arm, and that’s the platform’s largest line item, full stop. Cost scales with examples times arms times repeats, and the part that doesn’t parallelize is curation — a human has to verify each one, and that time doesn’t compress no matter how much compute you throw at it. The inverse number is the one that should sit on every dashboard: at n=500 your set can’t see anything smaller than about six points, at n=1000 it’s four and a half. Report the accuracy without that number and you’ve invited the exact error this whole architecture exists to prevent.
Host: And then there’s the part nobody puts in the architecture diagram — who’s allowed to touch the dataset itself.
Guest: Right, because that dataset is a leak surface — it’s curated, high-value, often real customer data sitting in a system built for convenience, not confidentiality. Worse, if it ever reaches a training corpus, contamination gives you no warning, just scores that quietly improve. So access to change an expected answer is access to change the verdict, which means edit access needs review, not just write permission for whoever’s running an experiment — and if you’re using a judge grader, remember it’s an injection surface too, since a graded output that can address the judge can influence its own score. Watch four things going forward and you’ll catch most of what we’ve covered today before it costs you: smallest detectable delta, undecidable rate, provenance coverage, and grader disagreement.
Not covered
The planner wanted these and found nothing in the source to support them:
- Vendor-specific pricing or throughput benchmarks for running LLM-as-judge graders at scale
- Guidance on choosing an embedding model for a semantic cache used alongside this evaluation platform
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Problem
Section titled “Problem”Every change to an AI system — a new model, an edited prompt, a different retrieval strategy, a routing policy — arrives with a claim that it is better. An evaluation platform exists to test that claim, and most of them do not, because they are built to produce a number rather than to support a comparison.
The gap between those two jobs is where the real failures live. A harness that reports “84% → 88%” has produced two numbers and no decision. Whether that four-point movement is an improvement, noise, or a disguised regression depends on how many examples produced it — and at the dataset sizes teams actually use, the honest answer is usually that nobody can tell.
Three things have to hold before an evaluation number means anything, and each fails silently on its own:
- The grader has to be able to fail. A harness that scores everything highly is indistinguishable from a system that works.
- The dataset has to be able to resolve the difference being claimed. Below a certain size, a set cannot see a small improvement no matter how real it is.
- The dataset has to measure correctness, not agreement with the system’s own past output.
Requirements
Section titled “Requirements”- A grader whose failure mode is tested, with a deliberately broken system asserted to score zero and a perfect one asserted to score 100%.
- Significance testing on every comparison, reported as an interval rather than a point estimate.
- A published smallest detectable delta for each evaluation set, so a claim can be checked against what the set can actually see.
- Dataset provenance recorded per example — who verified this answer, and against what.
- Instability named rather than averaged, so flaky examples are visible as flaky.
- A path to the model that bypasses caches, or the evaluation measures infrastructure instead of the system.
Constraints
Section titled “Constraints”- Effect size enters the sample-size calculation squared. Halving the delta you want to detect more than quadruples the examples required. This is arithmetic, not a tooling limitation, and it is the constraint that makes most evaluation sets too small.
- Small sets break the approximation itself. Below roughly five expected outcomes per cell, a normal approximation is not merely imprecise — it is inapplicable, which is a different statement and demands a different answer.
- Graders are systems too. A
containsgrader passes an output that states the right answer and its opposite, which is what verbose model output looks like. An LLM-as-judge grader is the least verifiable kind and needs its own evaluation before its verdicts mean anything. - Datasets grown from production output drift toward measuring consistency. An example whose expected answer came from the system’s earlier output cannot detect a regression that was always there, and nothing in the resulting accuracy number reveals it.
- Comparing many variants multiplies the chance of a false positive. Five prompt variants against one baseline is five chances to get lucky, and correcting for that pushes the required sample size up again.
Request Flow
Section titled “Request Flow”flowchart TD
Base["Baseline run<br/>42 / 50 correct"] --> Cmp
Cand["Candidate run<br/>44 / 50 correct"] --> Cmp
Cmp["Comparison<br/>delta = +4.0 points"] --> Usable{"Expected cell counts >= 5?<br/>is the approximation valid here?"}
Usable -->|No| Unknown["is_significant() -> None<br/>'this set cannot tell'<br/>NOT 'no difference'"]
Usable -->|Yes| CI["95% confidence interval<br/>on the delta"]
CI --> Zero{"Does the interval contain zero?"}
Zero -->|"Yes — [-9.6%, +17.6%]"| NotReal["Not established.<br/>Same data is consistent<br/>with a 10-point regression."]
Zero -->|No| Real["Established.<br/>Report the interval, not the point."]
NotReal --> Fix["Compute what the set can see:<br/>n=50 at 85% -> 20 points minimum<br/>to see +3 points, need 2,033 per arm"]The path is a pipeline with one unusual property: it is allowed to refuse to answer.
Run both arms over the same dataset, then grade each output against its expected answer. Both arms seeing identical examples is what makes a paired test available later, and paired tests need fewer examples for the same power.
Compare with an interval, not a difference. The point estimate is the least informative output of the whole pipeline, and it is the one that gets quoted.
Return one of three verdicts. Significant, not significant, or undecidable — the approximation is unusable at this size. That third state is the design decision that distinguishes an evaluation platform from a scoreboard, because collapsing it into “not significant” reports no difference when the truth is this set cannot tell. Those get acted on very differently: the first closes the investigation, and the second should open one into the eval set.
Failure Modes
Section titled “Failure Modes”A harness that cannot fail
If nothing in the suite asserts that a deliberately broken system scores zero, every other number the platform produces is unverified. This is the meta-test, and it is the one most evaluation code does not have — which means a grader silently returning “correct” on an exception looks exactly like a system that works. The same shape appears one layer out, in the pipeline that runs the harness: see Module 12 on why a green build is only evidence once you can show which checks actually ran.
Reporting a point estimate
“84% → 88% on 50 examples” is the shape of most improvement claims. The 95% interval on that measurement runs from −9.6% to +17.6%: the same data is consistent with a ten-point regression. The point estimate is not a measurement, and quoting it without the interval is how a regression ships as an improvement.
A dataset that cannot see the claimed improvement
At an 85% baseline, fifty examples cannot resolve anything below roughly twenty percentage points. Every smaller movement such a set reports is noise being read as signal — and it will report them consistently enough to build a quarter’s worth of roadmap on.
Expected answers taken from the system's own output
The set now measures agreement with a former self. It cannot detect a regression that was present when the examples were captured, it rewards consistency over correctness, and the accuracy number looks entirely normal. Provenance has to be a field, because it is invisible in the result.
Substring graders on verbose output
A contains check scores an answer as correct when the output contains the right answer somewhere,
including when it also contains the opposite. Longer model outputs make this more likely over time,
so grader precision degrades quietly as models get chattier.
Averaging across repeats
Running each example several times and averaging hides instability behind a plausible number. An example that passes 50% of the time and one that always passes at 50% partial credit are different problems; the average makes them identical.
Evaluating through a cache
The suite measures the cache, not the system, and will keep reporting the previous model’s scores after an upgrade — which presents as a change that did nothing. See Semantic Response Caching.
Scaling
Section titled “Scaling”- Cost scales with examples × arms × repeats, and the sample-size table below is a direct statement of how expensive a small claimed improvement is to verify. This is the number that decides an evaluation platform’s budget, and it is usually discovered after the platform is built.
- Runs are embarrassingly parallel across examples, so wall-clock is a concurrency and rate-limit question rather than an architectural one. The upstream provider’s limits are the real ceiling.
- Paired designs buy power for free. Both arms already run the same examples, so a paired test (McNemar’s) is more powerful than the unpaired calculation and requires fewer examples for the same confidence — worth doing, and worth checking whether it changes any decision, because it often does not change the order of magnitude.
- Results storage grows with every run, and its value is longitudinal: the ability to ask whether a regression appeared three releases ago is the main thing a platform offers over a script.
- Dataset curation is the bottleneck that does not parallelise, because every verified example costs human attention, and that is precisely what the sample-size tables demand more of.
Security
Section titled “Security”- Evaluation datasets are a leak surface. They are curated, high-value, and often contain real customer data captured as examples, held in a system built for convenience rather than confidentiality.
- Contamination is the integrity risk. An evaluation set that reaches a training or fine-tuning corpus stops measuring anything, and there is no signal when it happens — only scores that improve.
- Access to change expected answers is access to change the verdict, so the dataset deserves review on edit rather than write access for anyone running an experiment.
- A judge model is an injection surface. If a graded output can address the judge, the system under test can influence its own score.
Trade-offs
Section titled “Trade-offs”Dataset size vs. what you are allowed to claim
A small set is cheap, fast, and can only support large claims. A large set costs curation effort proportional to the precision you want, and precision costs quadratically. The trade is not “bigger is better” — it is choosing the smallest delta worth acting on and then paying for exactly that, which is a decision most teams have never explicitly made.
Exact graders vs. judge models
Exact and structural graders are verifiable, cheap, deterministic, and only work where the answer has a checkable form. A judge model grades anything and is the least verifiable component in the system — it needs its own evaluation, against its own verified set, before its verdicts carry weight. The backing lab omits a judge grader entirely for that reason, which is a defensible scoping decision and a real limitation.
Unpaired vs. paired tests
The unpaired two-proportion test is simpler, more conservative, and easier to explain to a room. A paired test exploits the fact that both arms saw the same examples and needs fewer of them. The honest reason to reach for paired is power; the honest reason to stay unpaired is that a conservative number that nobody disputes is worth something, and the difference rarely changes the order of magnitude.
Offline evaluation vs. online measurement
Offline evaluation is repeatable, safe, and measures a proxy. Online measurement — A/B on real traffic — measures the thing you care about and costs exposure of real users to the worse arm, plus the time to accumulate significance at your traffic level. The sample-size arithmetic is identical in both; only the currency changes, from curated examples to user sessions.
The sample sizes are the cost, and stating them plainly is most of what this architecture is for.
Examples needed per arm to detect an improvement at 95% confidence and 80% power:
| Baseline | +1pt | +2pt | +3pt | +5pt | +10pt |
|---|---|---|---|---|---|
| 70% | 32,644 | 8,077 | 3,551 | 1,248 | 291 |
| 80% | 24,638 | 6,036 | 2,626 | 903 | 197 |
| 85% | 19,458 | 4,722 | 2,033 | 683 | 138 |
| 90% | 13,493 | 3,211 | 1,353 | 432 | 71 |
The inverse is the number that belongs on the dashboard — the smallest delta a set can see, at an 85% baseline:
| n | Smallest visible delta |
|---|---|
| 30 | 25.8% |
| 50 | 20.0% |
| 100 | 14.1% |
| 500 | 6.3% |
| 1,000 | 4.5% |
| 5,000 | 2.0% |
Effect size enters squared, so halving the delta you want to see at least quadruples the dataset — measured at 4.95× rather than the naive 4×, because as the candidate rate approaches 1.0 its variance shrinks and the coarse measurement gets disproportionately cheap.
The other costs follow from these: model spend is examples × arms × repeats, and curation is human time proportional to the same number. A platform that wants to detect three-point improvements at an 85% baseline is committing to roughly two thousand verified examples per arm, and that commitment is the architecture’s largest line item.
Publish the smallest detectable delta next to the score
An accuracy number without its set’s resolution invites exactly the error the platform exists to prevent. Reporting “88% (n=50, smallest detectable delta 20pt)” makes a four-point improvement claim visibly unsupportable at the point where someone would otherwise act on it — which is cheaper than the alternative, where the claim becomes a roadmap item first and a retraction later.
Observability
Section titled “Observability”- Smallest detectable delta per evaluation set, published alongside every score it produces.
- Confidence intervals on every comparison, with the point estimate deliberately never shown alone.
- Undecidable-verdict rate. A rising share of comparisons the statistics cannot resolve is a dataset problem announcing itself before anyone acts on a bad number.
- Per-example stability across repeats, with unstable examples named rather than averaged, so they can be fixed or removed.
- Dataset provenance coverage — what fraction of examples have a human-verified expected answer versus one captured from system output. This is the metric that catches a set drifting toward measuring consistency.
- Grader disagreement rate, where more than one grader applies. Two graders diverging is the earliest available signal that one of them is wrong.
- Score history per example, not just per run, because a regression is usually a small set of examples flipping rather than a uniform decline.
Production Deployment
Section titled “Production Deployment”Before anyone makes a decision on its output
- A deliberately broken system is asserted to score 0%, and a perfect one 100%.
- Every comparison reports a confidence interval, and the point estimate never appears alone.
- Each evaluation set’s smallest detectable delta is computed and published with its scores.
- The significance test can return “undecidable”, and callers handle that distinctly from “not significant”.
- Every example records who verified its expected answer and against what.
- Examples captured from the system’s own output are flagged as such.
- Graders are tested against verbose output that contains both an answer and its opposite.
- Repeated runs name unstable examples rather than averaging over them.
- The evaluation path bypasses response caches.
- Multiple-variant comparisons apply a correction, and the corrected sample size is what the set is measured against.
- Dataset edits are reviewed, and the set is excluded from any training or fine-tuning corpus.
Hands-on Lab
A running implementation: graders, a two-proportion significance test that returns True, False,
or None, and the sample-size tables above computed rather than quoted. Includes the meta-test —
a deliberately broken system asserted to score zero — because an eval that cannot fail is not an
eval. No model is called; systems under test are deterministic functions with a configured
accuracy, and the statistics are unpaired, which makes the sample sizes conservative.
Read the lab documentation →
labs/evaluation-platformproduction-shaped
Interview Questions
Section titled “Interview Questions”A teammate reports 84% → 88% on 50 examples. What do you say?
That the interval on that measurement runs from about −9.6% to +17.6%, so the same data is consistent with a ten-point regression. Fifty examples at that baseline cannot resolve anything below roughly twenty percentage points. The right next question is not whether the change is good but what the eval set’s smallest detectable delta is — and if nobody knows, that is the finding.
You want to detect a three-point improvement at an 85% baseline. What does that cost?
Roughly two thousand examples per arm at 95% confidence and 80% power. The important part is why the number is so large: effect size enters squared, so halving the delta you want to see more than quadruples the dataset. That arithmetic is what makes most evaluation sets unable to support the claims made from them, and it is a curation cost in human time, not a compute cost.
Why should a significance test be allowed to return neither true nor false?
Because “not significant” means we found no difference and small sets need to say this set
cannot tell, which is a different fact with a different response. Below about five expected
outcomes per cell the normal approximation is inapplicable rather than imprecise. Collapsing that
into False closes an investigation that should have been opened — into the dataset rather than
the change.
What is the first test you write for an evaluation harness?
That a deliberately broken system scores 0%, and its twin, that a perfect one scores 100%. Every other number the harness produces is unverified until those pass — a grader that silently returns “correct” on an exception is indistinguishable from a system that works. An eval that cannot fail is not an eval.
Your eval set was built from production traffic and the system's own answers. What is wrong with it?
It measures agreement with a former self rather than correctness. It cannot detect a regression that was already present when the examples were captured, and it rewards consistency — including consistent errors. Nothing in the accuracy number exposes this, which is why provenance has to be a field on each example rather than an assumption about the set.
You test five prompt variants against one baseline and one wins by four points. Do you ship it?
Not on that evidence. Five comparisons is five chances to get lucky, so the threshold has to be corrected for multiplicity — which raises the sample size the set needed in the first place. Ship it only if the winner clears the corrected bar on a set whose smallest detectable delta is below four points, and treat a result that clears an uncorrected threshold alone as a hypothesis to re-test, not a finding.