Architecture: Task-Based Model Routing
Read the transcript
1. The static config and the backwards intuition
Host: So picture the setup almost every team ends up with: you’ve got more than one model available, and somewhere at the top of a config file, someone made a decision about which one handles what. Once. And then that decision just quietly runs on every single request forever.
Guest: Right, and it fails in one of two directions. Either you’ve got a strong, expensive model answering questions a five-cent model would have nailed, or you’ve got the cheap model getting handed work it was never built to do, and just failing at it. Both are the same underlying mistake, just pointed opposite ways.
Host: Okay, so the obvious fix seems easy — just sort your models by price and default to the cheapest one that’s available. What’s wrong with that?
Guest: Nothing, until you remember price has zero relationship to capability. That gives you a router that’s fast and cheap and confidently wrong. And even once you fix that — even once you’ve got a system that genuinely checks capability first — there’s a second trap waiting, one that sounds completely reasonable: escalation, try the cheap model first, fall back to the expensive one when it looks unsure. Everybody’s intuition says that’s obviously smart. Measured the right way, it’s actually backwards.
2. What a real router has to guarantee
Host: Okay, before we get into that escalation trap, let’s nail down what a router actually has to promise. If you had to write the non-negotiables on a whiteboard, what’s on the list?
Guest: Six things, and they’re in order for a reason. Capability filtering before cost, because a model that can’t do the task isn’t a cheap option, it’s not an option at all. Then a quality floor with a visible outcome when nothing in the fleet clears it — no silent shrug. A budget cap that doesn’t silently degrade quality, deterministic selection so the same task and fleet always produce the same choice, an auditable reason attached to every decision rather than just a model name, and last, cost accounting the router doesn’t get to flatter itself with — including the case where its own strategy turns out to be the expensive one.
3. Trust, decay, and the corner nobody accounts for
Host: Okay, so you’ve got six ordered checks, all deterministic and auditable. That sounds like it should be airtight. What breaks it?
Guest: Trust, mostly. When a model gets registered as capable of reasoning, the router just believes it — there’s no way to verify that claim, only to make it explicit so a failure is attributable instead of mysterious. And quality scores are the same problem in slow motion: they’re a snapshot of a moving target, so the router can quietly optimize against numbers that stopped being true weeks ago, and nothing in the system throws an error when that happens.
Host: So it can be confidently wrong and still look like it’s working. What about the confidence signals themselves — can those degrade too?
Guest: Worse — you can swap an informative confidence signal for pure noise and nothing catches it. No type checker, no code review flags it, escalation still fires, still bills twice, and just returns less. And people forget escalation has a latency cost too — a task that escalates pays two round trips, and if there was a deadline, the cheap-then-expensive path can miss it when the single expensive call up front wouldn’t have.
4. Why the ordering is the design: capability, then policy, then escalation
Host: So let’s zoom out from the failure modes for a second and just look at the shape of the thing. If I’m reading the request flow right, the order itself is basically the argument — capability first, then policy, then escalation. Why does that ordering matter so much that you’d call it the design rather than just a diagram?
Guest: Because capability is a hard filter, not a preference. Everything downstream is an optimization over whatever survives that filter, so if you run cost before capability you get exactly the failure we already talked about — picking the cheapest model that literally cannot do the task. That’s not a tuning problem, it’s a one-line ordering mistake, and it’s the whole reason the flow is written the way it is.
Host: Okay, so capability filters the field. Then policy — floor and budget — runs over survivors. What happens when policy itself can’t be satisfied, when nothing left clears the floor or fits the budget?
Guest: Both have to fail loudly, but differently. A floor nothing clears still returns a decision — you can’t serve nothing to the user — but it records that the floor was missed, so it’s a logged failure, not a silent one. A budget cap is stricter: it raises an error instead of quietly downgrading to something worse. And escalation sits outside all of this — it’s not a routing rule, it’s a second call, accounted for separately, including in the billing. The architecture treats it as its own event because pretending it’s just another branch in the router is how people lose track of what it actually costs.
5. The measurement that broke the assumption it was built to confirm
Host: So you ran the numbers on escalation itself, not just the plumbing around it, and I gather the test suite you wrote to prove it saves money actually failed. What happened?
Guest: We gave it the best possible case — a perfect confidence signal, low exactly when the answer’s wrong, which no real system gets. Over two hundred tasks, escalation pushed correct answers from 166 to 197, a 19% accuracy gain. But cost per correct answer went from 0.00024 to 0.00279 — eleven times worse.
Host: Eleven times worse while accuracy went up? That sounds like it should be a contradiction, not just a bad tradeoff.
Guest: It feels like one, but it’s arithmetic, not a modelling failure. The cheap model is already right 83% of the time, so it’s generating a huge pile of cheap correct answers. Escalation only adds a small number of expensive correct answers on top, so the average can only go up — even when the escalation policy is working perfectly. Cost per correct answer just isn’t built to notice that a policy is succeeding; it punishes any expensive addition to an already-cheap base, which is why it’s the metric everyone reaches for and the one that’s lying to them.
6. The number that actually answers the question
Host: So if cost per correct answer is lying to us, what’s the number that actually tells the truth here?
Guest: Marginal cost per rescued answer. Instead of averaging over everything, you isolate what escalation specifically bought: the escalated cost minus the baseline cost, divided by the escalated correct minus the baseline correct. That ratio is just the price of the answers escalation actually fixed, nothing else mixed in.
Host: And when you actually run that math, what does it look like?
Guest: At a 2x price ratio between the cheap and strong model, you’re paying about 0.00044 per rescued answer. Push the ratio out to 75x and it’s about 0.016. Both are real numbers you can defend, but neither one tells you if it’s worth paying — that’s a separate question about what a wrong answer costs your product, and the router has no opinion on that, it just hands you the honest bill.
7. How this fails quietly in production
Host: So let’s get concrete about how this actually breaks, because I don’t think it breaks loudly. Walk me through the failure modes one at a time — where does it go wrong first?
Guest: First one: cheapest model wins on price before anyone checks if it can do the job. Small is the cheapest thing in the fleet and it cannot write code, so a router that sorts by price first hands it code tasks anyway — fast, cheap, confidently wrong. And here’s the quiet part: a budget cap that silently downgrades instead of raising an error is a quality regression that shows up in zero dashboards. It reads as a cost win. Then there’s the floor nobody clears — if you set a quality bar above everything in the fleet and the router just shrugs and returns something anyway, that floor is decorative, unless you force it to record the miss in the reason string, which is the only thing that makes an unsatisfiable bar visible instead of silently ignored.
Host: And the tie-breaking one — that’s the one that sounds small but corrupts your data, right?
Guest: Exactly, if two candidates are equal and the router picks based on dict iteration order or registration sequence, your cost reports don’t reconcile run to run, and any A/B test you run on top of that is just noise dressed up as a result. But the worst one is the confidence signal going bad, because nothing in the code, the types, or a review catches it — escalation still fires, you still get billed twice, you just stop rescuing almost anything. It’s identical on paper and completely different in reality, which is exactly why you have to measure it instead of trusting that it compiles.
8. Escalate or go straight to the strong model — and who gets to decide
Host: So let’s put escalation next to its blunt alternative: just send everything straight to the strong model. Why would anyone ever pick the two-hop route over the simpler one-hop route?
Guest: Because escalation buys the same accuracy for two round trips and two bills, while going straight there buys it for one of each but at a higher baseline cost every single time. The choice really turns on how often the cheap model is already right and what a wrong answer actually costs you. Most teams that adopt escalation have never written either number down, they just assumed the two-hop path was cheaper because it feels cheaper.
Host: And that’s the same trap as before — cost per correct answer looks fine, but it’s not the number that tells you whether escalation is pulling its weight.
Guest: Right — and once you’ve settled on the number that actually matters, there’s a separate question of who gets to act on it. A router that owns the escalation decision can enforce floors, caps, and capability filtering globally, but it frustrates the caller who already knows their task needs the strong model. Letting callers pin a model directly is responsive and hands them the spend dial, but then you lose that central control. The version that keeps both properties is letting callers express requirements — capabilities, a quality floor — while the router still picks the actual model.
9. Locking down who can spend your money
Host: So if callers only express requirements and never name a model, where does the actual security risk hide in that registry? It feels like the interesting failure mode isn’t the caller anymore.
Guest: Right, it moves upstream to whoever writes the capability entries. A model that claims a capability it doesn’t have gets routed work it will fail, so that registry has to be reviewed like code, not edited like a config file. And separately, budget caps need to be per tenant, not global, or one runaway caller eats everyone’s ceiling — that’s abuse control as much as cost control. Then when something does go wrong, the routing reason for that specific request has to be sitting in the audit log, because which model served it and why is the very first question anyone asks once a bad answer reaches a customer.
10. What scales, what doesn’t, and what to watch
Host: Let’s talk scale, because I’d assumed the router itself was the expensive part — all that capability filtering and policy checking on every request. Is that actually a bottleneck?
Guest: No, and that’s worth saying plainly: the routing decision is basically free. It’s a filter and a sort over a small in-memory registry, microseconds against a model call that takes seconds — nothing about that needs to be distributed. If your router has turned into a network hop, that better be justified by wanting one central place for policy, not by some performance need, because performance was never the problem. What does scale unpredictably is the load on the expensive tier, because the escalation rate is just a function of how accurate the cheap model is on live traffic right now, and if that drifts down even a little, you’re quietly shoving a growing share of everything onto the smallest, priciest tier you have.
Host: So the router stays cheap, but the consequence of it making slightly worse calls doesn’t stay small — it lands entirely on your most constrained capacity. How do you actually see that happening before the bill or the queue tells you?
Guest: You watch escalation rate and escalation yield together, never rate alone — rate just tells you spend, yield tells you whether the escalations are still rescuing anything, and a rising rate with falling yield is exactly the degrading-signal pattern, invisible in either number by itself. Alongside that you want quality-floor miss rate, which tells you whether the fleet can’t meet the bar or the bar’s wrong, and model mix drift, which is the cleanest tell of all — traffic hasn’t changed but more of it is landing on the expensive model, meaning the cheap model’s real-world accuracy has moved out from under you.
11. Shipping it, and the evaluation pipeline it leans on
Host: So if someone’s shipping this tomorrow, what actually has to be in place before it goes out the door? Give me the checklist version.
Guest: Three things, non-negotiable. The ordering is enforced by a test, not a comment, so a cheaper incapable model literally cannot be selected even by accident. Escalation gets its own line in the cost record, and that policy stays justified over time rather than just once. And the model spec’s quality field has a named human owner with a review cadence, because a number nobody’s accountable for just drifts until it’s wrong. But none of that holds up on its own — the router only makes good decisions if the quality scores feeding it stay honest over time, and keeping them honest as models and prompts and traffic all shift underneath you is a whole separate discipline, continuous model evaluation, with its own hard problem of building a grader that can actually fail and a dataset big enough to see the difference you’re claiming. The router assumes that problem is solved elsewhere. It has to be — that’s this whole episode in one sentence.
Not covered
The planner wanted these and found nothing in the source to support them:
- A concrete production case study of a company adopting this router and the before/after business impact
- Details of how the confidence signal used for escalation is itself trained or calibrated
- A worked example of setting a numeric value on ‘what a wrong answer costs’ for a specific business
- Specific latency numbers (milliseconds) for an escalated double round trip versus a single strong-model call
- Comparison of this architecture against named third-party routing frameworks or products
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Problem
Section titled “Problem”A platform with more than one model has to decide, per request, which one serves it. Most systems make that decision once, statically, at the top of a config file — and then pay for it on every request forever, in one of two directions: a strong model answering questions a cheap one would have got right, or a cheap model failing at work it was never capable of.
The obvious fix is to sort the fleet by price and take the cheapest. That produces a router which is cheap, fast, and wrong, because price ordering has nothing to do with whether a model can do the task at all.
The less obvious problem is the one that survives a competent design. Once capability is handled, teams reach for escalation — try the cheap model, fall back to the expensive one when confidence is low — on the reasoning that it captures most of the quality at a fraction of the cost. That reasoning is widespread, intuitive, and, measured against the metric everybody uses to justify it, backwards.
This architecture is about making the decision explicit, ordered, auditable, and measured against a number that does not lie.
Requirements
Section titled “Requirements”- Capability filtering before cost. A model that cannot do the task is not a cheap option; it is not an option.
- A quality floor, and a visible outcome when nothing in the fleet clears it.
- A budget cap that does not silently degrade quality.
- Deterministic selection. The same task and the same fleet produce the same choice, run to run.
- An auditable reason attached to every routing decision, not just the model name.
- Cost accounting the router does not flatter — including the case where the router’s own strategy is the expensive one.
Constraints
Section titled “Constraints”- Capability declarations are trust, not verification. A model registered as able to reason will be routed reasoning work whether or not it can do it. The router has no way to check the claim; it can only make the claim explicit and the failure attributable.
- Quality scores decay. They are static claims about a moving target, and a router optimising against stale quality numbers degrades without any component reporting an error.
- Escalation spends latency as well as money. An escalated task pays two round trips and can miss a deadline the single expensive call would have met — the corner of the cost-latency-quality triangle that routing discussions most often drop.
- The router cannot know what a wrong answer costs. That number is a product decision. A router that assumes one has embedded a business judgement in infrastructure, where nobody will find it.
- Confidence signals vary in usefulness, invisibly. Swapping an informative signal for a meaningless one changes nothing a type checker or a code review would catch. Escalation still fires, still bills twice, and returns far less.
Request Flow
Section titled “Request Flow”flowchart TD
Task["Task<br/>task_class · estimated_tokens"] --> Cap{"Which models declare<br/>support for this task class?"}
Cap -->|None| Err["NoCapableModel<br/>raised, not guessed"]
Cap -->|"Capable set, cheapest first"| Pol{"Policy"}
Pol -->|CheapestCapable| P1["First in the capable set"]
Pol -->|QualityFloor| P2["Cheapest clearing the floor<br/>or best available, reason recorded"]
Pol -->|BudgetCapped| P3{"Above the ceiling?"}
P3 -->|Yes| Budget["BudgetExceeded<br/>raised, never a silent downgrade"]
P3 -->|No| P1
P1 --> Run["Run on the chosen model"]
P2 --> Run
Run --> Conf{"Confidence below threshold?"}
Conf -->|No| Done["Answer · one call billed"]
Conf -->|"Yes, and a better model exists"| Esc["Run best model too<br/>BOTH calls billed"]
Conf -->|"Yes, already on the best"| Done
Esc --> Marginal["Answer · two calls billed<br/>the rescued answer has a price"]The ordering carries the design:
Capability first, because it is a hard filter and everything after it is an optimisation over whatever survives. Running cost before capability is what produces the cheapest-model-that-cannot-do-it failure, and it is a one-line mistake.
Policy second — quality floor, then budget cap — over the models that remain. Both can fail to be satisfiable, and both must fail loudly: a floor nothing clears still returns a decision, because you cannot serve nothing, but records that the floor was missed. A budget cap raises an error rather than quietly picking something worse.
Escalation last, and separately accounted. It is not a routing rule but a second call, and the architecture treats it as one — including in the billing.
Failure Modes
Section titled “Failure Modes”Sorting by price before filtering by capability
The cheapest model in a fleet is frequently the least capable one. A router that sorts first and filters second returns it, quickly and confidently, for tasks it cannot perform. The system looks efficient in every cost dashboard and is failing the work.
A budget cap that silently downgrades
Choosing a weaker model to stay under budget is a quality regression that raises no alert, appears in no error rate, and shows up in the cost report as a success. Raising an error is the only version where someone finds out.
A quality floor nobody clears, invisibly
If the floor is set above anything in the fleet and the router silently ignores it, the floor is decorative — it constrains nothing and nobody knows. Recording the miss in the decision’s reason makes an unsatisfiable bar auditable instead.
Ties broken by iteration order
A router whose choice between equal candidates depends on dictionary ordering or registration sequence produces cost reports that cannot be reconciled between runs, and A/B results that are noise. Break ties on a stable key.
Escalation on an uninformative confidence signal
Escalation is only worth anything if the signal fires when the answer is actually wrong. A meaningless signal produces the identical code path, the identical double billing, and almost none of the rescued answers — and nothing in the type system, the API, or a code review distinguishes the two cases. It is measurable and nothing else.
Optimising cost per correct answer
The trap metric, and the one everybody reaches for. See below: it moves the wrong way for arithmetic reasons that have nothing to do with whether escalation is working.
Scaling
Section titled “Scaling”- The routing decision itself is free. It is a filter and a sort over a small in-memory registry, microseconds against a model call measured in seconds. Nothing here needs to be distributed, and a router that has become a network hop should be justified by policy centralisation, not performance.
- The model registry is a hot read that changes rarely, which makes it cacheable and makes cache invalidation the mechanism by which a withdrawn model actually stops receiving traffic.
- Escalation multiplies load on the expensive tier non-linearly. The escalation rate is a function of the cheap model’s accuracy on live traffic, so a distribution shift that degrades the cheap model quietly redirects a growing share of the workload onto the tier with the smallest capacity and the highest price.
- Per-tenant policy multiplies the decision space, not the cost. Floors, caps, and fleets differ per tenant; the evaluation stays local, so this scales with configuration size rather than traffic.
Security
Section titled “Security”- Treat capability declarations as a trust boundary. A model that claims a capability it lacks gets routed work it will fail, so the registry is a place where a wrong entry has consequences and should be reviewed like code, not configuration.
- Do not let callers name the model. A caller-specified model is a direct spend lever pointed at your account, and the obvious abuse — request the most expensive tier for everything — needs no sophistication at all. Callers describe the task; the router picks.
- Budget caps are an abuse control as well as a cost control, and they belong per tenant rather than globally, where one caller’s runaway consumes everyone’s ceiling.
- The routing reason belongs in the audit record. Which model served a request, and why, is the first thing anyone asks after a bad answer reaches a customer.
Trade-offs
Section titled “Trade-offs”Escalation vs. going straight to the strong model
Escalation buys accuracy for two round trips and two bills. Going straight to the strong model buys the same accuracy for one of each, at higher baseline cost. The choice turns entirely on how often the cheap model is right and what a wrong answer costs — and the honest version of this trade requires both numbers, which most teams adopting escalation have never written down.
Cost per correct answer vs. marginal cost per rescued answer
Cost per correct answer is the intuitive metric and it is misleading whenever the cheap model is already mostly right. Marginal cost per rescued answer measures what escalation actually buys. The first is easier to compute and easier to present; the second is the one that supports a decision.
Static quality scores vs. continuously measured ones
Static scores are simple, reproducible, and wrong by a growing margin. Measured scores track reality and require an evaluation pipeline, labelled traffic, and a feedback path — which is another system to build and operate (see Continuous Model Evaluation). The intermediate position — static scores with an owner and a review cadence — is worth more than it sounds, because the common failure is not staleness but staleness nobody is responsible for.
Router-owned policy vs. caller-specified model
A router that owns the decision can enforce floors, caps, and capability filtering globally, and frustrates the caller who knows their task needs the strong model. Letting callers pin a model is responsive and hands them the spend dial. Letting callers express requirements — capabilities, a quality floor — while the router picks the model keeps both properties.
This is where the architecture earns its keep, and where the measurement contradicted the design it was written to validate.
Escalation improved accuracy 19% and made cost per correct answer 11× worse. Over a fixed 200-task workload with a perfect confidence signal — low exactly when the answer is wrong, the best case any escalation policy could achieve:
| Cheap only | Escalating | |
|---|---|---|
| Correct | 166 / 200 | 197 / 200 |
| Total cost | 0.040 | 0.550 |
| Cost per correct answer | 0.00024 | 0.00279 |
The relationship is not an artefact of the price gap between those tiers. It holds at 75×, 10×, 5×, and 2×.
The cause is arithmetic, not modelling. A cheap model already right 83% of the time contributes a large number of cheap correct answers. Escalation adds a smaller number of expensive ones, so the average cost per correct answer can only move upward — even when escalation is working perfectly. That is what makes it a trap: the metric does not distinguish a policy that is failing from one that is succeeding.
What escalation actually buys is marginal correct answers, so the number that supports a decision is what each rescued answer cost:
marginal cost per rescued answer = (escalated_cost − baseline_cost) / (escalated_correct − baseline_correct)About 0.00044 per rescued answer at a 2× price ratio, and about 0.016 at 75×. Whether either is worth paying depends on what a wrong answer costs — which is a product question, and one the router should not pretend to answer.
The remaining cost levers:
- Capability filtering is the largest single saving, because it is what stops the strong model serving requests the cheap one handles correctly.
- Budget caps bound the worst case per tenant, which matters more than the average when one caller can move the bill.
- Escalation latency is a cost paid in the deadline budget, not the invoice, and it is the one that surfaces as a user-visible timeout rather than a line item.
Observability
Section titled “Observability”- Decisions by model, with the reason string. Not just which model served the traffic, but which rule selected it — capability filter, floor, cap, or default.
- Escalation rate and escalation yield, together. The rate alone says how much you are spending; the yield — rescued answers per escalation — says whether the confidence signal is worth anything. A rising rate with a falling yield is the signature of a degrading signal, and it is invisible in either metric alone.
- Marginal cost per rescued answer, tracked over time, so the escalation policy is continuously justified rather than justified once at adoption.
- Quality-floor miss rate, which distinguishes a fleet that cannot meet the bar from a bar set wrong.
- Budget-cap rejections per tenant, which is a capacity conversation, not an error.
- Model mix drift. A gradual shift toward the expensive tier with no change in traffic is the clearest available signal that the cheap model’s real-world accuracy has moved.
Production Deployment
Section titled “Production Deployment”Before real traffic
- Capability filtering runs before any cost comparison, with a test that a cheaper incapable model is never selected.
- The quality floor has a defined, visible outcome when nothing in the fleet clears it.
- The budget cap raises rather than silently selecting a weaker model.
- Tie-breaking is on a stable key, not iteration order, and is asserted in a test.
- Callers cannot name a model directly; they describe the task or its requirements.
- Escalation is separately accounted, and both calls appear in the cost record.
- The escalation policy is justified by marginal cost per rescued answer, and that number is monitored rather than computed once.
- Every decision records the model and the reason in a form the audit path keeps.
ModelSpec.qualityhas a named owner and a review cadence.- Per-tenant budget caps exist, so one caller cannot exhaust a shared ceiling.
Hands-on Lab
A running implementation: capability filtering, quality floors, budget caps, deterministic
tie-breaking, and the escalation measurement above — including
test_marginal_cost_per_rescued_answer_scales_with_the_cost_ratio, which pins the relationship so
the conclusion cannot quietly invert. No model is ever called; correctness comes from a seeded hash
so the numbers are reproducible, and latency is not modelled at all.
Read the lab documentation →
labs/model-routerproduction-shaped
Interview Questions
Section titled “Interview Questions”You have four models at different prices. How does a request pick one?
Filter by capability first, then apply policy — quality floor, then budget cap — over what survives. The ordering is the answer: sorting by price before filtering by capability selects the cheapest model that cannot do the task, which is fast, cheap, and wrong. Cost is an optimisation over the capable set, never a filter on the full fleet.
Your team wants to adopt escalation to cut costs. What do you measure first?
Marginal cost per rescued answer, not cost per correct answer. Cost per correct answer moves the wrong way whenever the cheap model is already mostly right — a model correct 83% of the time contributes many cheap correct answers, and escalation adds fewer expensive ones, so the average rises even when escalation works perfectly. In one measurement it rose 11× while accuracy improved 19%. The metric that supports the decision is what each additional correct answer cost, weighed against what a wrong answer is worth.
How would you know your escalation confidence signal had stopped being useful?
Escalation yield — rescued answers per escalation — tracked alongside escalation rate. A useless signal produces the same code path, the same double billing, and far fewer rescued answers, and nothing in the API, the types, or a code review distinguishes it from a good one. Rate alone tells you what you are spending; only the yield tells you what you are buying.
Why shouldn't the budget cap just pick a cheaper model?
Because that is a quality regression disguised as a cost success. It raises no error, appears in no error rate, and shows up in the cost report as the system working. Raising instead makes the trade-off someone’s explicit decision rather than an invisible default — the cap has told you the budget and the required quality are incompatible, which is information.
Should callers be allowed to specify the model?
No — a caller-specified model is a spend lever pointed at your account, and the abuse requires no sophistication. Let callers express requirements instead: the capabilities the task needs and the quality floor it demands. The router keeps capability filtering, floors, and caps enforceable centrally, and the caller still gets the strong model when the task genuinely requires it.
Where does a static quality score go wrong?
It decays. It is a fixed claim about a fleet that changes underneath it — new versions, changed defaults, shifting traffic — and a router optimising against a stale score degrades with no component reporting an error. The strong fix is measured quality from an evaluation pipeline; the cheap and underrated one is a named owner and a review cadence, because the usual failure is not that the number is stale but that nobody is responsible for it.