Skip to content

Architecture: Model Serving Platform

Listen to this page12:43
Read the transcript

1. The trap: throughput problem wearing a latency budget

Host: So let’s start with the trap at the center of this whole episode. Everyone treats model serving like it’s a latency problem — get the response back fast — but you’re telling me that’s the wrong frame entirely.

Guest: Right, it’s actually a throughput problem wearing a latency budget as a disguise. The accelerator you’re running on is brutally expensive, and it’s only earning its keep when it’s crunching many requests at once in a batch. But every single request in that batch has to wait for the batch to fill before processing starts, so the very thing that makes the hardware efficient is the thing that makes any individual request slow.

Host: And you can’t just pick a lane — tune for one and the other collapses. So what makes this harder than normal capacity planning, where you’d just add more servers when things get busy?

2. What a serving platform must guarantee — and what it can’t wish away

Guest: Because normal capacity planning assumes the trade-off is solvable — you throw resources at it and both latency and utilization improve together. Here they’re structurally opposed: batching more raises utilization but adds queuing delay, and there’s no dial that wins both. You have to pick a matched point for your actual traffic and SLO, and that’s before you even count the other guarantees a serving platform has to hold — safe rollout to a small slice of traffic, rollback without a redeploy, p95 and p99 visibility since the mean hides exactly the tail damage batching introduces, and isolation so one noisy model doesn’t starve another on shared hardware.

Host: That’s a long list to hold at once. Which of those is the one that actually bites hardest in practice — where does the design get genuinely stuck rather than just inconvenienced?

Guest: Two things don’t bend. First, scaling has to be driven by a leading signal, not a trailing one — if a replica takes three minutes to come up, any trigger with less than three minutes of lead time is already too late, no matter how good your autoscaler logic is. Second, accelerator memory caps how many requests you can serve concurrently, and it’s usually not the model weights that eat that budget — it’s the KV cache growing with every sequence in flight, which quietly sets a hard ceiling under whatever throughput number you were hoping for.

3. Inside the batcher: racing size against timeout

Host: So walk me through what actually happens when a request lands. You said scaling has to lead demand — but before it even gets to a replica, it’s sitting in some queue, right?

Guest: Right, it joins a pending queue for whichever version it was routed to, and a background loop closes that batch the moment either the size limit is hit or the wait timer expires — whichever fires first. That race is the whole design: it means no request waits longer than the timeout even during a lull, but a burst of traffic still fills batches to the size limit instead of leaving throughput on the table. And critically, each model version runs its own batcher with its own metrics, so a canary can race those two triggers on completely different settings than stable without the two ever sharing a queue.

Host: And that’s not just a nice-to-have — that’s why you’d never trust a size-only or timeout-only trigger alone?

Guest: Exactly, each one fails at an opposite end of the traffic range. Size-only maximizes throughput but lets a single lonely request wait forever when traffic is thin; timeout-only bounds latency but wastes capacity under load because it closes batches half-empty. Racing both gets you the better half of each, at the price of two coupled parameters you now have to tune together against a real traffic distribution instead of picking either one in isolation.

4. Where it actually breaks in production

Host: So you tune the batcher, you’ve got both knobs racing against each other — where does that actually go wrong once it’s live? Give me the real incident, not the theory.

Guest: Start with latency itself lying to you. A request that arrives just after a batch closes waits nearly the full timeout longer than one that slips in just before — that’s a bimodal distribution, and mean latency is the arithmetic average, so it just erases the whole story. You have to look at p95 and p99, because those are the only numbers that show what your batching config is actually costing someone.

Host: That’s already unsettling for something as basic as a dashboard. What about the canary — I’d assume that’s the safety net that catches the bad stuff before it spreads.

Guest: It should be, but if it’s only taking five percent of traffic and fails outright, that moves your blended error rate by maybe five points — which looks like normal noise. The rollout mechanism reports healthy while the exact failure it exists to catch is happening in production. Same pattern shows up in autoscaling — react to current load with replicas that take minutes to boot, and capacity shows up after the spike and gets clawed back before the next one, so the system just oscillates, paying for capacity it never has when it needs it.

5. Scaling on signals that lead demand, not trail it

Host: So if utilization is watching the wrong thing, what’s the right signal? You need something that tells you demand is rising before the replicas are underwater.

Guest: Queue depth and wait time — they lead demand instead of trailing it, because a queue starts growing the instant requests arrive faster than you can drain them, well before any GPU shows sustained high utilization. Pair that with keeping warm headroom instead of scaling to zero, since a cold start on the critical path usually costs more in tail latency than an idle replica costs in dollars. And once you’re managing capacity, batch size and replica count are two separate levers with different consequences — a bigger batch buys throughput at the cost of queueing delay, another replica buys throughput at the cost of money, and which one you pull depends on whether it’s latency or spend that’s actually under pressure.

Host: So it’s not one autoscaling knob, it’s two levers you’re choosing between based on which SLO is screaming. What happens when you also start stacking multiple live versions on top of that for canarying?

6. Safe rollout and rollback that’s actually fast

Guest: Two separate mechanisms, actually. Shadow evaluation runs the candidate on mirrored traffic and throws away the output — it exposes nobody, which costs double compute and proves nothing about whether the model is actually good, because nobody real ever saw the answer. Canarying is the traffic split that follows: you route a real slice of users to the candidate and watch quality signals on live behavior, so the usual sequence is shadow first for correctness, canary second for quality.

Host: Okay, so say the canary goes bad at 2am. How fast can you actually get back?

Guest: That’s the split that matters most under pressure. Weight-based rollback just flips routing back to a version still loaded in memory — that’s seconds, no build step, which is the only reason it’s usable mid-incident, but it means the old version has to stay resident, and that memory cost caps how many versions you can keep live at once. Redeploy-based rollback frees that memory but means rebuilding and reloading, which is far too slow when you’re bleeding traffic — and underneath both, version pinning has to actually hold, because a caller entitled to a specific version for compliance reasons can’t be silently rerouted just because a canary decision changed something upstream.

7. Data and weights are both assets to protect

Host: Let’s pivot to security, because I think people picture that as a perimeter problem — who can call the API — and miss that the serving path itself is full of leakage points. Prompts and completions are user data flowing through queues, batchers, logs, traces. Where does that actually go wrong?

Guest: The classic failure is sampled request logging done after the fact — you log the raw prompt for debugging, and now sensitive content is sitting in a system with completely different retention rules than what your data policy promised. Redaction has to happen before logging, not as a cleanup pass afterward, because after is always too late for whatever already got shipped to a log aggregator. There’s a batching-specific version too: a shared batch is a shared failure domain, so if a crash or an error affects one request in the batch, you need to be sure batch composition never crossed a tenant isolation boundary that actually matters. And separately from data, the weights themselves are an asset — anyone with access to the serving host has access to the weights, which is a supply-chain and IP problem that lives right alongside the data protection one, not instead of it.

8. The cost model and the dashboard that must reflect it

Host: So walk me through the money side, because I think people hear ‘batching’ and think latency, not cost. Where’s the actual spend hiding?

Guest: Idle accelerator time is the dominant line item — that’s the thing that makes batching a cost mechanism before it’s a performance one, since every empty slot in a batch is money burning. Padding waste stacks on top of that, because mixed-length sequences mean you’re paying compute for the padding tokens too, so length-aware batching is really a cost optimization wearing a latency costume. And it’s not just active traffic — every live version costs standing memory whether it’s serving anything or not, so your canary and rollback headroom is a permanent line item, not something that only shows up during an incident.

Host: So the dashboard has to make all of that legible, not just latency. What actually needs to be on it?

Guest: Percentiles split into queue wait versus model time, because mean latency actively lies to you in a batching system — and everything per-version, error rate, latency, throughput, quality, never blended, or the canary can’t do its job. Then batch size distribution and time-to-fill, because if batches keep closing on timeout instead of hitting size, that wait is pure added latency you’re paying for with no throughput benefit. And you need accelerator utilization with KV cache split from weights, plus some crude quality signal — refusal rate, output length, whatever you’ve got — because infrastructure health can look perfect while the actual output is quietly getting worse.

9. The pre-traffic checklist, and a lab that makes the mechanism tangible

Host: So if I’m about to flip traffic onto this thing, what’s actually on the checklist before I let it near production? Walk me through it as a gut check, not a spec.

Guest: Batch size and wait time tuned against your real traffic length distribution, and re-tuned on a schedule because that distribution drifts. p95 and p99 alerted, never the mean. Every version’s metrics tracked independently so a canary’s failure is provably visible, not blended away. Rollback measured as an actual wall-clock number, not assumed fast. Autoscaling driven by queue depth or wait time with warm headroom sized from a measured cold start, not a guess. A quality signal on the serving path that’s distinct from health checks, redaction applied before anything gets logged, and multi-model packing decisions that record which models share a tier so you know your blast radius. If you can’t check off all nine, you don’t know what you’re serving, you’re hoping.

Host: That’s a good place to land — that’s basically the whole episode compressed into nine bullets. For anyone who wants to feel this instead of just hear it, there’s a lab that runs the batching race and the canary routing live.

Guest: Right, it’s a real batcher racing size against timeout, one queue per version, promote and rollback implemented as routing-weight changes, and a p50/p95/p99 harness so you can watch the tail behave the way we described. One honest caveat: the model call is a fixed-cost sleep, so the throughput curve looks perfectly linear — double the batch, halve the cost, forever, which real inference never does because padding makes short sequences pay for the longest one in their batch. The lab’s upfront about that limit, which is exactly why it ships a sweep exercise instead of a recommended number — the right batch size is a property of your workload’s length distribution, not something a demo can hand you.

Not covered

The planner wanted these and found nothing in the source to support them:

  • A deep dive into prefill/decode mechanics and continuous batching internals, which the source only gestures at via a link to a separate module
  • The SLO and burn-rate alerting machinery referenced as a sibling lab, since no excerpt here details how it decides a canary is failing
  • Line-by-line walkthrough of the lab’s test suite or the ASGI lifespan bug, which is implementation trivia rather than architectural argument

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

Inference is a throughput problem wearing a latency budget. The accelerator is expensive and idle unless it is processing many requests at once; every request in a batch waits for that batch to form. Optimizing either objective alone produces a system that fails on the other — a batch-size-only tuning gives excellent utilization and unusable tail latency, and a latency-only tuning gives an accelerator running at a fraction of what it cost.

Two further properties make this unlike ordinary service capacity planning. Cold start is measured in minutes, not seconds, because weights have to be loaded — so reactive autoscaling adds capacity after the spike that triggered it has passed. And a new model version can be worse without being broken: faster, well-formed, and wrong. No health check detects that.

  • Bounded added latency. Batching must not let any request wait longer than a stated bound.
  • High accelerator utilization, since idle capacity is the dominant cost.
  • Safe version rollout. A new model reaches a small traffic slice first, with its own metrics.
  • Fast rollback, without a deploy.
  • Tail-latency visibility. p95 and p99, because the mean hides the failure batching introduces.
  • Cold-start-aware scaling, driven by a signal that leads demand rather than trailing it.
  • Workload isolation, so one model or tenant cannot starve another on shared hardware.
  • The throughput/latency trade cannot be optimized away, only positioned. There is no configuration that wins on both; there is only one matched to the actual traffic and SLO.
  • Cold start bounds how reactive scaling can be. If a replica takes three minutes to serve traffic, any signal with less than three minutes of lead time is too late by construction.
  • Batch cost is not linear in batch size. It grows with the padded sequence length of the batch, so one long sequence makes every short one in that batch pay for its length.
  • Accelerator memory caps concurrency, and the KV cache — not the weights alone — is usually what actually caps it.
  • Quality regressions are invisible to infrastructure. Liveness, readiness, and error rate all stay green while output quality falls.

A request joins the pending queue for its routed version and awaits its own result. A background loop closes the batch when either the size limit is reached or the wait timer expires — whichever fires first — so no request waits longer than the configured bound even during a lull, and high traffic still fills batches.

Each model version owns its own batcher and its own metrics. That is what lets a canary run different batch tuning from the stable version without the two coupling through a shared queue, and what keeps a canary’s error rate from being averaged into the aggregate.

Tail latency nobody is watching

A request arriving just after a batch closes waits nearly the full timeout longer than one arriving just before. Mean latency hides this completely — it is the arithmetic average of a bimodal distribution. Only p95 and p99 show the cost the batching configuration is actually imposing.

Canary metrics folded into the aggregate

A canary taking 5% of traffic and failing outright moves a blended error rate by five points — easily inside normal variance. The rollout mechanism reports healthy while the thing it exists to protect against is happening. Per-version metrics are not a nice-to-have; without them the canary is theatre.

Cold-start-blind autoscaling

Scaling on current utilization when replicas take minutes to become ready means capacity arrives after the spike and is then scaled back down before the next one. The system oscillates, paying for capacity that is never available when needed.

Batch size tuned once, on the wrong distribution

A batch configuration tuned on short prompts behaves differently the moment long ones appear, because padding waste grows with the spread of sequence lengths in a batch. Traffic mix drifts continuously; a launch-day tuning is usually wrong within months and nothing signals it.

Health checks that pass while quality regresses

A new version that returns fast, well-formed, wrong answers passes every infrastructure check. Readiness, liveness, latency, and error rate are all green. Only an evaluation signal — offline scores, or an online proxy — distinguishes it from a good version.

Noisy-neighbour accelerator sharing

Packing several models onto one device raises utilization and couples their latency. One model’s traffic spike degrades every other model on that device, and the symptom appears in a service whose own traffic did not change — which is the hardest kind of incident to diagnose.

  • Scale on queue depth and wait time, not utilization. Both lead demand; utilization trails it, and with minutes of cold start, trailing signals are useless.
  • Keep warm headroom rather than scaling to zero, unless the traffic pattern is genuinely scheduled. The idle cost of a warm replica is usually smaller than the tail-latency cost of a cold start on the critical path.
  • Batch size and replica count are alternative capacity levers with different latency consequences: a larger batch adds queueing delay, another replica adds cost. The right mix depends on which side of the SLO is under pressure.
  • Length-aware batching beats naive batching at scale, because grouping similar-length sequences reduces the padding waste that makes large batches stop paying for themselves.
  • Version-count multiplies memory pressure, since each live version holds its own weights and KV cache. Canarying is not free in capacity terms, which bounds how many versions can be live.
  • Prompts and completions are user data, frequently sensitive, and pass through queues, batchers, logs, and traces. Every one of those is a place they can be retained by accident.
  • A shared batch is a shared failure domain. A crash while processing a batch affects every request in it, so batch composition must not cross a tenant isolation boundary that matters.
  • Model weights are assets. Access to the serving host is access to the weights, which is a supply-chain and IP concern distinct from data protection.
  • Version pinning must be enforceable. A caller entitled to a specific version for compliance or reproducibility reasons must not be silently served another by a routing change.
  • Redact before logging, not after. Sampled request logging is the standard way prompt content escapes into systems with different retention rules.

Racing a size trigger against a timeout

A size-only trigger maximizes throughput and lets a lone request wait indefinitely at low traffic. A timeout-only trigger caps latency and wastes capacity under load. Racing both takes the better half of each, at the cost of two coupled parameters that must be tuned together against a real traffic distribution rather than independently.

Canary by traffic split vs. shadow evaluation

A traffic split exposes real users to a candidate version and yields real quality signal. Shadow evaluation runs the candidate on mirrored traffic with output discarded, exposing nobody — and costing double compute while proving nothing about user-visible behaviour. Shadow first for correctness, canary for quality, is the usual sequence.

Dedicated accelerators vs. multi-model packing

Dedicated capacity gives predictable latency and leaves expensive hardware idle between spikes. Packing raises utilization and couples the latency of unrelated models. The deciding question is whether the models share an availability tier — packing a latency-sensitive model with a batch workload is the specific mistake.

Rollback by routing weight vs. by redeploy

Weight-based rollback is seconds and needs no build, which is what makes it usable during an incident. It requires the previous version to still be loaded, consuming memory that bounds how many versions can be live. Redeploy-based rollback frees that memory and is far too slow when it matters.

  • Idle accelerator time is the dominant line item, which is what makes batching a cost mechanism before it is a performance one.
  • Padding waste is real spend. Batching mixed-length sequences means paying for the padding, so length-aware batching is a cost optimization as much as a latency one.
  • Every live version costs memory whether or not it serves traffic, so canary and rollback headroom is a standing cost, not an incident-time one.
  • Warm headroom is bought deliberately as insurance against cold-start latency, and should be sized from the traffic pattern rather than left at whatever the default was.
  • Tokens are the unit that scales with traffic, so context growth upstream — in RAG, in agent loops — lands here as a serving bill.
  • p50, p95, and p99 latency, split into queue wait and model time. Mean latency is actively misleading for a batching system.
  • Per-version everything — error rate, latency, throughput, and a quality signal — never blended, or the canary mechanism cannot function.
  • Batch size distribution and time-to-fill. Batches consistently closing on timeout rather than size mean the size limit is not the binding constraint and the wait is pure added latency.
  • Queue depth and wait time as the scaling signal, exported at a resolution useful for a controller rather than a dashboard.
  • Accelerator utilization and memory headroom, with KV cache occupancy separated from weights, since that is what actually caps concurrency.
  • A quality metric on the serving path, however crude — refusal rate, output length distribution, downstream acceptance. Infrastructure health cannot substitute for it.

Before real traffic

  • Batch size and wait time were tuned against the real traffic length distribution, and the tuning is dated and re-run on a schedule.
  • p95 and p99 are alerted on; the mean is not used as the latency SLI.
  • Every version’s metrics are tracked independently, and a canary’s failure is provably visible in them.
  • Rollback is a routing-weight change, and its wall-clock duration has been measured.
  • Autoscaling is driven by queue depth or wait time, with warm headroom sized from cold-start duration.
  • Cold-start time is measured, not estimated.
  • A quality signal exists on the serving path, distinct from health checks.
  • Prompt and completion redaction is applied before anything is logged or traced.
  • Multi-model packing decisions record which models share an availability tier.

Hands-on Lab

A running implementation of the batching and rollout mechanics: a batch closing on whichever of size or timeout fires first, canary routing with independent per-version metrics, promote and rollback as routing changes, and a p50/p95/p99 benchmark harness. The model call is a fixed-cost sleep, which the lab is explicit about hiding the padding-waste curve. Read the lab documentation →

labs/dynamic-batching-inferenceproduction-shaped

Why does a batch close on whichever of size or timeout fires first?

Because each trigger alone fails at one end of the traffic range. Size-only lets a single request wait indefinitely during a lull; timeout-only caps latency but leaves throughput on the table under load. Racing both bounds the worst-case wait at the timeout while still filling batches when traffic allows. The cost is two parameters that must be tuned together against real traffic.

You doubled the batch size and throughput barely moved. Why?

Padding waste. Batch cost grows with the padded sequence length, so a batch containing one long sequence makes every short sequence in it pay for that length. Past a point, adding requests adds padding rather than useful work. Optimal batch size is therefore a function of the traffic’s length distribution, not its rate — which is why length-aware batching exists and why a launch-day tuning goes stale.

Why can't you autoscale inference on CPU or GPU utilization?

Because cold start is minutes, and utilization is a trailing signal. By the time utilization is high, the spike is already being served badly, and a replica started now becomes ready after it has passed. Queue depth and queue wait time lead demand, which is what a controller needs when the actuator is slow. Warm headroom covers the remainder.

Your canary is at 5% and the dashboard is green. What would you not have seen?

A canary failing outright moves a blended error rate by five points, which is inside normal variance for most services. If metrics are not tracked per version, the rollout mechanism reports healthy exactly when it should be firing. Independent per-version metrics are what make a canary a control rather than a gesture.

How do you detect that a new model version is worse, not broken?

Not with health checks — a worse model returns fast, well-formed, wrong answers and passes liveness, readiness, latency, and error rate. It needs a quality signal: offline evaluation before rollout, and an online proxy during it, such as refusal rate, output length distribution, or a downstream acceptance metric. Infrastructure health and output quality are independent axes.

Why is rollback a routing change rather than a deploy?

Speed, during the one situation where speed matters most. A weight change is seconds and needs no build; a redeploy is minutes under incident pressure. The cost is that the previous version stays loaded and consumes memory, which bounds how many versions can be live at once — a standing capacity cost paid for incident-time recovery speed.