Dynamic Batching Inference
Read the transcript
1. The core mechanism: racing size against timeout
Host: So today we’re digging into dynamic batching for inference serving, and the whole thing hinges on one deceptively simple trick: a batch closes when either it fills up or a timeout hits, whichever comes first. Why not just pick one of those and call it a day?
Guest: Because each one alone breaks in an opposite way. If you only trigger on size, a lone request arriving during a quiet stretch just sits there waiting for company that may never show up, so your latency is unbounded. If you only trigger on a timeout, you’re safe on worst-case wait, but under heavy load you’re closing tiny batches every few milliseconds when you could’ve packed in way more work per GPU pass.
Host: So racing them means you get the throughput win when traffic is heavy and the size trigger fires first, but you still cap the pain when traffic is sparse and the clock runs out instead.
Guest: Exactly, and that cap is a real number you control, the max wait time in seconds — it’s the worst case any single request can be stuck waiting on strangers to fill the batch. That’s the core mechanism the whole lab builds on, and everything else we’ll talk about, canaries, per-version metrics, rollback, all sits on top of this one race.
2. Canary rollout as routing, not deploys
Host: So once you’ve got that race dialed in, how do you actually roll out a change to it without risking the whole fleet? I know the lab spins up two versions side by side.
Guest: Right, stable-v1 and canary-v2, each with its own batcher, its own queue, its own weight — nine to one in the demo. That isolation matters because canary-v2 can run a totally different batch size and wait time without touching stable’s tuning, and critically, the two never share a queue. But the part people get wrong is the metrics: if you blend stable and canary traffic into one aggregate error rate, a canary that’s failing badly on its small slice just gets averaged away into noise, and you’ve defeated the entire point of doing a gradual rollout in the first place.
Host: So the whole safety story depends on watching canary-v2’s numbers in isolation, not the combined feed. And when something does go wrong, what does pulling it back actually look like?
Guest: That’s the other piece — promote and rollback are just weight changes, hitting an endpoint to shift traffic, not redeploying anything. That’s the difference between rollback taking a second during an incident versus you standing there redeploying the old version while requests keep failing. Cheap rollback is what makes the canary trustworthy, because you’ll actually use it under pressure instead of hesitating.
3. Why the mean lies and the fixed-cost sleep hides the real curve
Host: So if rollback is the safety net, how do you even know when you need it — like what’s the actual signal that batching is hurting you rather than helping?
Guest: You can’t see it in the mean, which is exactly why the lab benchmarks p50, p95, and p99 under bounded concurrency instead of just averaging. A request that lands right after a batch closes sits there for almost the full timeout, while one that lands right before gets swept in immediately — the mean smooths that gap away, but the tail is where the actual user experience lives. If you’re only watching the mean, you’ll ship a config that looks fine and quietly punishes a chunk of your traffic every cycle.
Host: Okay, so the tail tells the truth about latency. But you’ve also said the lab’s cost model has a real blind spot — walk me through that.
Guest: The handler just sleeps a fixed duration no matter how many payloads are in the batch, so doubling batch size always halves per-request cost — a clean line that goes forever. Real inference doesn’t work that way: cost tracks padded sequence length, so one long sequence in a batch drags every short sequence along to pay for it, and past some point a bigger batch stops paying for itself. That breakeven point is set by the traffic’s length distribution, not its rate, which is exactly why the lab hands you a sweep harness instead of a recommended number — the mechanism and the measurement transfer, the curve doesn’t.
4. A factory, not a singleton — and the bug that proved why
Host: So we’ve got the mechanism, the canary as routing, the honest curve — what’s the last piece? You mentioned something about the app itself being built differently than you’d expect.
Guest: Right, create_app builds a fresh router and batchers every time you call it, instead of one shared instance for the whole process — because a DynamicBatcher can’t restart once shutdown has run, same as a real batching worker once its background task dies. That’s not theoretical: httpx’s ASGITransport doesn’t drive the lifespan protocol, so in early tests the batcher’s background loop never started, and every request to infer just awaited a batch that nothing would ever flush. The suite didn’t fail, it hung — which is its own lesson, that a missing lifespan doesn’t throw an error, it just quietly starves you, and the fix was making the app a factory and making tests enter the lifespan explicitly rather than assuming it’s there.
Not covered
The planner wanted these and found nothing in the source to support them:
- Comparing this lab’s batching mechanism directly against vLLM’s continuous batching and PagedAttention implementation details
- Discussing the SLO-driven autoscaling or async gateway labs as alternatives to this lab’s rollout mechanism
- Security and multi-tenancy concerns around prompt/completion data in the batching queue
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Serving inference is a throughput problem wearing a latency budget: batching amortizes an expensive model call across many requests, but every request in a batch waits for the batch to form. This lab implements the mechanism that bounds that wait, plus the canary routing and percentile benchmarking needed to change a model version without betting all traffic on it.
Source: labs/dynamic-batching-inference
Batch formation
Section titled “Batch formation”flowchart TB
R1["Request A"] --> Q
R2["Request B"] --> Q
R3["Request C"] --> Q
Q["Pending queue<br/>each request awaits its own future"]
Q --> RACE{"Batch closes on<br/>whichever fires first"}
RACE -->|"max_batch_size reached"| F["Flush"]
RACE -->|"max_wait_seconds elapsed"| F
F --> H["handler(payloads) → outputs<br/>one model call for the whole batch"]
H --> RES["Each pending future resolved"]
subgraph Router["CanaryRouter: one batcher per version"]
direction LR
SV["stable-v1<br/>weight 9"]
CV["canary-v2<br/>weight 1"]
end
Router -->|"independent batch tuning,<br/>independent metrics"| QWhat it demonstrates
Section titled “What it demonstrates”- Racing a size trigger against a timeout. A size-only trigger lets a lone request wait
indefinitely during a lull; a timeout-only trigger caps latency but wastes throughput under load.
Racing both gets the better half of each, and bounds the worst-case wait at
max_wait_seconds. - One batcher per model version. A canary can run its own batch size and wait time without perturbing the stable version’s tuning, because the two never share a queue.
- Per-version metrics kept separate. An aggregate error rate blending stable and canary traffic can look flat while the canary alone is failing badly — which defeats the entire purpose of a gradual rollout.
- Promote and rollback as routing changes. Both are weight adjustments, not deploys, which is what makes rollback fast enough to actually use during an incident.
- Percentile benchmarking with bounded concurrency. p50/p95/p99, because the mean hides exactly the failure batching introduces: a request arriving just after a batch closes waits nearly a full timeout longer than one arriving just before.
What the fixed-cost stand-in hides
Section titled “What the fixed-cost stand-in hides”The handler sleeps for a fixed duration regardless of how many payloads are in the batch. That makes the throughput win look linear in batch size — double the batch, halve the per-request cost, forever.
Real inference does not behave that way. Cost grows with the padded sequence length of the batch, so a batch containing one long sequence makes every short sequence in it pay for that length. Past some point a larger batch stops paying for itself, and where that point falls depends on the traffic’s length distribution, not just its rate.
This is why the lab ships a benchmark harness and a sweep exercise rather than a recommended batch size: the right configuration is a property of the workload, and this lab’s simulated cost model cannot produce a curve you should trust. What it can teach is the mechanism and the measurement method — which transfer intact.
Run it
Section titled “Run it”cd labs/dynamic-batching-inferencepython3.12 -m venv .venvsource .venv/bin/activatepip install -e '.[dev]'uvicorn inference.app:app --reloadTwo demo versions start: stable-v1 (weight 9, ~30ms simulated batch cost) and canary-v2
(weight 1, ~15ms — an optimized candidate on a small traffic slice).
# Weighted routing across versionscurl -s -X POST localhost:8000/v1/infer \ -H 'content-type: application/json' -d '{"payload": 42}'
# Per-version weights and metrics, tracked independentlycurl -s localhost:8000/v1/versions
# Throughput and tail latencycurl -s -X POST localhost:8000/v1/benchmark \ -H 'content-type: application/json' -d '{"num_requests": 200, "concurrency": 20}'
# Promote and roll back as routing changes, not deployscurl -s -X POST localhost:8000/v1/versions/canary-v2/promotecurl -s -X POST localhost:8000/v1/versions/canary-v2/rollbackVerify it
Section titled “Verify it”pytest # 27 testsruff check .mypy srcWhy the app is a factory, not a singleton
Section titled “Why the app is a factory, not a singleton”create_app() builds a fresh router and batchers per call rather than exposing one process-wide
instance. A DynamicBatcher cannot be restarted once shutdown() has run — the same constraint a
real batching worker has once its background task exits.
This surfaced as a genuine bug during development: httpx.ASGITransport does not drive the ASGI
lifespan protocol, so the batcher’s background loop never started and every request to /v1/infer
awaited a batch nothing would ever flush. The test suite hung rather than failed. The fix was both a
factory and tests that enter the lifespan explicitly — worth knowing for any FastAPI service whose
background workers are started in lifespan.
Principal-level discussion points
Section titled “Principal-level discussion points”- Batching trades a bounded latency increase for a large throughput gain by amortizing a fixed cost.
max_batch_sizeandmax_wait_secondsare one knob pair with opposing effects, and neither is correct independent of the traffic pattern and SLO. - Canary weighting only functions as a safety mechanism if per-version metrics are independent. A canary’s errors hidden inside an aggregate defeat the rollout entirely.
- Promote and rollback are cheap because they change routing, not infrastructure — compare a deploy-based rollback, which requires redeploying the previous version under incident pressure.
- A benchmark reporting only the mean hides the specific failure batching introduces. p95 and p99 are not optional here; they are the metric the mechanism is traded against.
- Optimal batch size is a function of the workload’s sequence-length distribution, not just its request rate — which is why a configuration tuned at launch is often wrong a year later.
Related
Section titled “Related”- Architecture: Model Serving Platform — the design-review companion: problem framing, constraints, cost, observability, and a pre-traffic checklist.
- Module 9: Model Serving — prefill/decode, continuous batching, and KV cache management this lab’s mechanism sits underneath.
labs/slo-driven-ai-operations— the SLO and burn-rate machinery that would decide when a canary is failing.labs/async-ai-gateway— theproduction-readyreference this lab is measured against.