Module 9: Model Serving
Read the transcript
1. The two workloads hiding inside one request
Host: So we’re kicking off the model serving module, and I want to start with something that trips up a lot of engineers: serving an LLM is not just ‘run inference, but faster.’ There are actually two completely different workloads hiding inside every single request you send to a model.
Guest: Right, and they pull in opposite directions. When you process the prompt — what we call prefill — every token is already known, so the GPU can chew through all of them in parallel. It’s compute-bound, and it saturates the hardware beautifully. But then generation, decode, is the opposite: you’re producing one token at a time, each one depending on the last, so there’s no parallelism across the sequence. Every single step has to read the entire model’s weights and the whole KV cache from memory just to spit out one token, so you’re bottlenecked on memory bandwidth, not compute.
Host: Which is the counterintuitive part — a GPU that looks idle during decode isn’t actually idle at all, it’s fully busy just shuffling memory around for one token. And that one fact is basically the thesis for this whole module: continuous batching, KV cache tricks, quantization, parallelism strategies — every one of them exists to reconcile that split, so by the end you should be able to trace any serving cost or latency number straight back to prefill versus decode.
2. Why decode leaves the GPU idle-looking but fully busy
Host: Okay, so walk me through what’s actually happening on the chip during one of those decode steps. Because ‘idle but busy’ still sounds like a contradiction to me.
Guest: It’s not idle, it’s just not computing much — it’s moving data. To produce a single token, the GPU has to pull the entire model’s weights plus that request’s whole KV cache out of memory, and it does that same full pull for every single token you generate. That’s a massive memory transfer to produce one number, so the GPU sits there saturated on bandwidth while its compute units mostly wait around with nothing to chew on.
Host: So the compute units are starved, not because there’s no work, but because the work is dwarfed by the time spent fetching everything from memory. Which means the fix can’t be a faster GPU, that memory traffic is fixed per token — the only lever is making that same expensive memory pass produce tokens for more than one request at once.
3. Static batching’s hostage problem, and continuous batching’s fix
Host: Okay, so if the fix is batching more requests into that same memory pass, why not just batch them the obvious way — collect a batch, run them all together, return results together? What breaks there?
Guest: Because requests don’t finish at the same time. If you lock a batch together and one sequence generates a short answer while another keeps rambling on, that short one’s slot just sits there burning GPU cycles doing nothing until the whole batch is done. You’ve paid for capacity you’re not using, and worse, any new request has to wait in line for that entire batch to close out before it even starts.
Host: And that’s the p99 killer, right — someone’s simple request gets stuck behind a stranger’s essay, purely by bad luck of which batch they landed in.
Guest: Exactly, it has nothing to do with that request’s own size, which makes it maddening to debug. Continuous batching fixes it by re-checking admission at every single decode step — the moment a slot frees up, a queued request’s prefill drops right in, so the batch stays full and no one’s held hostage by someone else’s long generation.
4. The KV cache: why memory, not model size, caps concurrency
Host: So once the batch is staying full, what’s actually filling up the GPU’s memory while all this is happening? Because I’ve heard people say the model weights themselves aren’t usually the bottleneck.
Guest: Right, it’s the KV cache. Every token needs to attend to all previous tokens, and recomputing those key and value projections from scratch at every single decode step would be enormously wasteful, so instead you cache them in GPU memory and just reuse them, only computing the new token’s projections each step. The catch is that cache grows linearly both with sequence length and with how many sequences you’re running concurrently, and it’s sitting in the exact same finite GPU memory pool as the model weights.
Host: So concurrency is capped by whichever fills up first, and it’s usually the cache, not the weights.
Guest: Exactly, and early implementations made it worse by allocating one contiguous block per sequence sized for the worst-case max length, so a short conversation still reserved space for its worst-case length, fragmenting memory badly. PagedAttention fixed that by borrowing the OS trick of paged virtual memory — small fixed-size blocks that don’t need to be contiguous, tracked through a block table per sequence — which is why engines like vLLM can pack meaningfully more concurrent requests into the same GPU.
5. PagedAttention: borrowing an OS trick to stop wasting memory
Host: So walk me through the mechanics a bit more. When you say fixed-size blocks tracked through a block table, what does that actually buy you over the old contiguous approach, concretely?
Guest: Think of it like OS paging: instead of demanding one giant contiguous slab of memory upfront, you allocate small blocks on demand as the sequence actually grows, and a lightweight table maps logical positions in the sequence to wherever those blocks physically live. So a short conversation just uses a few blocks and stops, it never reserves space for a max-length sequence it’ll never reach, and the scheduler can grab any free block anywhere in memory rather than needing one big open contiguous chunk.
Host: And that near-eliminates fragmentation because you’re never stuck with a bunch of small unusable gaps between reservations.
Guest: Exactly, and the compounding effect is what matters in production: less wasted memory per sequence means more sequences’ KV caches fit simultaneously, which is precisely the capacity constraint we just established as the real ceiling. That’s the core mechanism behind why modern serving engines fit meaningfully more concurrent sequences in the same GPU memory, and it’s a big part of why the paper that introduced it also popularized continuous batching as the default architecture for open-source LLM serving engines.
6. Inside a scheduler: reading the admission-control code
Host: Okay, let’s make this concrete, because I think people nod along to ‘admission control’ without picturing what actually runs. Walk me through what happens when this scheduler’s step function fires on a single tick.
Guest: Sure. First it looks at the queue and asks, for each waiting request in order, would admitting this one push total tokens in flight over the KV budget? If yes, it stops right there — it doesn’t skip ahead to a smaller request behind it, so ordering matters. If a request fits, it moves from queue into in_flight and that same tick it gets its one prefill pass, which just flips a flag and does nothing else that tick, before it starts decoding.
Host: And that’s the part that surprised me — prefill is its own tick, separate from decode, even though it’s the same loop. Then eviction is just as blunt: hit your max tokens, you’re deleted and your budget’s back on the table.
Guest: Right, and that delete is the whole point — it’s not cleanup happening later on some GC schedule, it’s synchronous, same tick, so the very next call to admit sees that freed budget immediately. That’s continuous batching’s mechanism laid completely bare: no tensor math, no CUDA kernels, just a budget check, a one-tick prefill, a decode increment, and an eviction — and a real engine bolts real forward passes and page-table accounting onto exactly this skeleton.
7. Quantization: the 70B decision nobody should make on vibes
Host: So the scheduler skeleton is bolted onto real memory now, but there’s still this lever nobody wants to pull without a good reason: quantization. Walk me through what actually changes when you drop a 70B model from FP16 to INT8 or INT4.
Guest: Weights and often the KV cache itself get stored in fewer bits, which shrinks memory footprint and, just as importantly, the memory bandwidth you need to move those weights every decode step — that’s a direct speedup, and the freed memory means more KV cache pages fit, so more concurrent sequences. FP16 gives you the best accuracy but the least capacity; INT8 roughly doubles concurrent requests for usually-small quality loss; INT4 pushes capacity further but starts risking visible degradation on tasks that need precise reasoning or exact output formats.
Host: That ‘usually small’ and ‘starts risking’ are doing a lot of work there — how does a team actually decide which level to ship on, instead of just trusting whatever the model vendor’s benchmark card says?
Guest: They don’t trust it, because a vendor’s aggregate benchmark averages across tasks and can completely hide a regression that only shows up on tasks needing precise reasoning or exact-format outputs. The only defensible move is running the same evaluation harness from Module 4 against FP16, INT8, and INT4 on representative production requests, and comparing task-by-task, not just the aggregate score. And how much degradation is tolerable at each level isn’t an engineering call at that point — it’s a product decision, because it’s trading measured accuracy against a very concrete, very measurable jump in concurrent-request capacity.
8. Splitting a model that doesn’t fit on one GPU
Host: So quantization gets you more room on a single GPU, but at some point the model just doesn’t fit no matter what precision you pick. Once you’re literally splitting weights across GPUs, what are the actual options and how do you pick?
Guest: Two real answers, and they trade off differently. Tensor parallelism splits the matrix math inside each layer across GPUs, so every single layer needs a synchronization step, which means you need fast interconnect like NVLink or the communication overhead eats your latency gains alive. Pipeline parallelism instead hands whole layers to different GPUs and streams batches through like an assembly line, so it tolerates weaker interconnect, but you pay for it with bubbles — idle GPU time while later stages sit waiting on earlier ones — unless your batch is big enough to keep every stage fed.
Host: So it’s fast-but-demanding versus tolerant-but-bubbly. Is this something you reach for instead of just adding more replicas, or alongside it?
Guest: Alongside, and it’s a completely separate axis — that’s the part people conflate. Horizontal scaling, more replicas behind a load balancer, is what you do once a model already fits on a GPU and you need more throughput. Sharding via tensor or pipeline parallelism is what you reach for only when a single model doesn’t fit on one GPU at all; once it fits, you go back to adding replicas, not more shards.
9. When it breaks: capacity gaps and cold starts
Host: So walk me through how this actually breaks in production, because I imagine none of these failures announce themselves as capacity problems. They probably show up looking like something else entirely.
Guest: Exactly, and that’s what makes them dangerous. If admission control isn’t checking real remaining KV budget, a burst of long-context requests can blow past available memory and force preemptions or rejections mid-flight — from the outside that looks like the model server just crashed, when really it’s a capacity-planning gap nobody instrumented. Separately, under static batching, one long generation holds an entire batch hostage, so requests that finished their own work seconds ago are still waiting — and that shows up as p99 latency spikes that have nothing to do with any individual request’s size.
Host: And that second failure mode, the mystery p99 — that’s exactly the kind of thing an on-call engineer would burn hours on, chasing the wrong request. What about scaling out to fix capacity, does that save you in the moment?
Guest: Not in the moment it’s needed, no. Scaling a GPU replica from zero means loading multi-gigabyte weights onto freshly provisioned hardware, and that can take tens of seconds to minutes — an eternity against the latency budget of whatever request triggered the scale-out in the first place. That’s why you don’t scale reactively on current load alone; you keep a warm pool sized for burst absorption and, where you can, scale ahead of known traffic patterns predictively.
10. Multi-tenant isolation and the cost of skipping the re-eval
Host: Warm pools and predictive scaling handle the capacity side, but scaling out doesn’t help if the fleet is shared across customers. What happens when multiple tenants are hitting the same GPUs?
Guest: Then isolation becomes a security property, not just a performance nicety. If one tenant sends a burst of huge prompts and there’s no boundary on KV cache allocation or batch composition, that tenant can starve everyone else’s latency budget without ever touching their data. In a poorly isolated implementation you can even leak timing information — the shape and duration of one tenant’s requests becomes inferable by another just from watching response latencies on the shared server.
Host: So isolation is the last capacity-planning gap. Is there an analogous silent gap on the model side — something that looks fine operationally but quietly breaks correctness?
Guest: Yes, and it’s the one people skip most casually: shipping a quantized or distilled model because it fits more requests per GPU, without re-running the safety evaluation harness. A general accuracy benchmark often won’t catch a safety regression, so the model can look fine on your dashboards while behaving differently on exactly the inputs that matter most.
11. Taking it into your hands: the batching lab and closing synthesis
Host: So if someone’s still not convinced that head-of-line blocking is a real cost and not just a theoretical worry, is there a way to actually see it rather than take our word for it?
Guest: That’s exactly what the dynamic batching lab is for. You extend the continuous scheduler we’ve been discussing with a static one that only admits new requests once every in-flight request finishes, throw both at a mixed workload of short and long generations, and measure how long the short requests take under each. The mechanism is real production-shaped code, canary routing, promote and rollback, a full p50/p95/p99 harness, so you’re not asserting the gap, you’re plotting it, and watching a short request sit behind a long one for however many extra seconds is a much better teacher than any diagram.
Host: That feels like the right place to leave people: go quantify it yourself instead of trusting the slide. And it loops back nicely to where we started this whole episode, because every one of those interview questions we’d hand a candidate, why batch decode steps at all, why static wastes capacity, why it’s KV cache memory and not compute that caps you, why tensor versus pipeline parallelism trades off interconnect for bubbles, all of it is just the prefill-is-compute-bound, decode-is-bandwidth-bound split wearing a different costume.
Guest: Right, that split is the one fact everything else in this module derives from, so if a candidate or a colleague can explain that clearly, they can rebuild the rest of the reasoning on the spot instead of memorizing it. That’s the actual bar: not knowing the answers, but knowing why the two workloads hiding in one request force every one of those answers to be what it is.
Not covered
The planner wanted these and found nothing in the source to support them:
- Specific dollar-cost or GPU-model benchmarks comparing serving engines
- Detailed methodology for what a safety evaluation harness should test on a quantized model
- Real-world outage postmortems from named companies running these serving stacks
- Performance comparisons between vLLM, TensorRT-LLM, and other named serving engines
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Executive Summary
Section titled “Executive Summary”Serving an LLM efficiently is not “run inference, but faster” — it’s reconciling two workloads with opposite performance characteristics inside a single request. Processing a prompt (prefill) is compute-bound and parallelizes trivially across tokens; generating a response one token at a time (decode) is memory-bandwidth-bound and barely uses the GPU’s compute at all. Every serving-layer technique this module covers — continuous batching, KV cache management, quantization, parallelism — exists to keep an expensive GPU busy despite decode’s fundamentally low utilization, and every serving cost and latency number a principal engineer is asked to explain traces back to this split.
Mental Model
Section titled “Mental Model”A request to a model server passes through two phases with different bottlenecks:
- Prefill processes the entire input prompt in one forward pass. All prompt tokens are known up front, so this parallelizes across tokens and saturates GPU compute — it’s throughput-friendly and fast per token processed.
- Decode generates output one token at a time, and each new token depends on the one before it — it cannot be parallelized across the sequence. Each decode step reads the entire model’s weights and the growing KV cache from GPU memory to produce a single token, so decode is bottlenecked on memory bandwidth, not compute. Most of the GPU’s compute capacity sits idle during decode.
A GPU serving one request at a time is not GPU-bound, it's memory-bandwidth-bound
This single fact explains why naive single-request serving is so expensive: an idle-looking GPU during decode is still fully occupied moving weights and KV cache through memory for one token at a time. The fix isn’t a faster GPU, it’s batching more requests’ decode steps together so the same memory traffic produces tokens for many requests at once — which is exactly what continuous batching does, and why it’s the highest-leverage serving optimization before reaching for anything else.
Architecture
Section titled “Architecture”flowchart TB
Req[Incoming requests] --> Sched["Continuous batching scheduler"]
subgraph GPU["Model server (single GPU / replica)"]
Sched --> Prefill["Prefill: process full prompt (compute-bound)"]
Prefill --> KV[(KV cache)]
KV --> Decode["Decode: generate one token at a time (memory-bandwidth-bound)"]
Decode --> Sched
Decode --> Out[Stream tokens to caller]
end
Sched -.new request joins mid-batch.-> Prefill
KV -.evicted under memory pressure.-> Evict[Preempt / reject]
subgraph Fleet["Autoscaled fleet"]
GPU
GPU2["Model server (replica 2)"]
GPUN["Model server (replica N)"]
end
LB[Load balancer] --> GPU
LB --> GPU2
LB --> GPUN
Req --> LBStatic batching — wait for a full batch of requests, run them together, return all results together — wastes GPU time whenever one request in the batch finishes generating before the others: the GPU sits there doing nothing useful for that slot until the whole batch completes. Continuous batching (the scheduler in this diagram) instead evaluates admission at every decode step, so a finished request’s slot is immediately replaced by a new request’s prefill, keeping the batch full and the GPU’s decode throughput high without ever forcing later arrivals to wait for an unrelated batch to finish first.
Deep Dive
Section titled “Deep Dive”The KV cache. Attention needs each new token to attend to every previous token in the sequence. Recomputing that from scratch at every decode step would be wildly wasteful, so the key and value projections for every previous token are cached in GPU memory (the “KV cache”) and reused across decode steps — only the new token’s projections need computing each step. The KV cache grows linearly with sequence length and with the number of concurrent sequences, and it competes with model weights for the same finite GPU memory: KV cache capacity, not model size alone, is usually what actually caps how many concurrent requests a GPU can serve.
PagedAttention. Early KV cache implementations allocated one contiguous memory block per sequence sized for its maximum possible length, fragmenting GPU memory badly (a short sequence still reserved worst-case-length space) and wasting a large fraction of available memory. PagedAttention borrows the operating-system idea of paged virtual memory: the KV cache is allocated in small fixed- size blocks that don’t need to be contiguous, referenced through a per-sequence block table. This essentially eliminates the fragmentation problem and is the core mechanism behind why modern serving engines (vLLM popularized it) fit meaningfully more concurrent sequences in the same GPU memory than naive implementations.
Continuous batching, mechanically. Instead of a fixed batch that runs to completion, the scheduler re-evaluates the batch composition at every decode step: any sequence that finished generating is evicted immediately, freeing its KV cache pages, and a queued request can be admitted and prefilled into the now-open slot — interleaved with ongoing decode steps for the rest of the batch. This is what actually keeps GPU utilization high under real traffic, where requests have wildly different output lengths and arrive at different times.
Quantization. Storing model weights (and sometimes the KV cache itself) in lower precision — INT8 or INT4 instead of FP16/BF16 — cuts memory footprint and memory-bandwidth pressure roughly in proportion, which directly speeds up the memory-bandwidth-bound decode phase and lets more of the KV cache fit in the same GPU memory. The cost is a small, usually acceptable, accuracy degradation; how much degradation is acceptable is a product decision, not a purely technical one, and should be measured against the same evaluation harness Module 4 covers, not assumed to be negligible.
Parallelism strategies. A model too large for one GPU’s memory needs to be split across GPUs. Tensor parallelism splits individual layers’ matrix operations across GPUs, requiring fast interconnect (NVLink) because every layer needs a synchronization step — it reduces per-token latency but adds communication overhead per layer. Pipeline parallelism instead assigns different layers to different GPUs and streams batches through them like a pipeline — better interconnect tolerance, but at the cost of pipeline bubbles (idle GPU time while later stages wait on earlier ones) unless the batch is large enough to keep every stage busy.
Research Note
The paper that introduced PagedAttention and popularized continuous batching as the default architecture for open-source LLM serving engines — worth reading directly for the memory fragmentation numbers that motivated it.
Source: Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention" (vLLM)
Implementation
Section titled “Implementation”A minimal continuous batching scheduler, showing the admission-control decision (does a new request fit in the remaining KV cache budget?) and the prefill/decode distinction at the core of every production serving engine, stripped of the actual tensor math:
Admission control and the prefill/decode split, without the model math
from __future__ import annotations
from dataclasses import dataclass
@dataclassclass Request: id: str prompt_tokens: int max_new_tokens: int generated_tokens: int = 0 prefilled: bool = False
class ContinuousBatchScheduler: def __init__(self, kv_cache_token_budget: int): self.kv_cache_token_budget = kv_cache_token_budget self.in_flight: dict[str, Request] = {} self.queue: list[Request] = []
def _used_tokens(self) -> int: return sum( r.prompt_tokens + r.generated_tokens for r in self.in_flight.values() )
def admit(self, request: Request) -> None: self.queue.append(request)
def step(self) -> list[str]: while self.queue: candidate = self.queue[0] projected = self._used_tokens() + candidate.prompt_tokens if projected > self.kv_cache_token_budget: break self.in_flight[candidate.id] = self.queue.pop(0)
finished: list[str] = [] for request in self.in_flight.values(): if not request.prefilled: request.prefilled = True continue request.generated_tokens += 1 if request.generated_tokens >= request.max_new_tokens: finished.append(request.id)
for request_id in finished: del self.in_flight[request_id] return finishedstep() runs once per decode tick: it admits queued requests only while they fit the remaining KV
cache budget (_used_tokens), gives every newly admitted request one prefill tick before it starts
decoding, advances every already-prefilled request by one generated token, and evicts anything that
just hit its length limit — freeing that budget for the next admit call’s candidates on the very
next tick. A real engine replaces the token-count budget with actual KV cache page accounting and
the prefill/decode ticks with real forward passes, but the admission-control shape is the same.
Production Example
Section titled “Production Example”A team serving a 70B-parameter model on a fixed GPU budget needs to pick a quantization level. FP16 gives the best accuracy but the fewest concurrent requests per GPU; INT8 roughly doubles concurrent capacity for a small, usually-acceptable quality loss on most tasks; INT4 pushes capacity further but risks visible quality degradation on tasks requiring precise reasoning or exact-format outputs. The only defensible way to choose is running the same evaluation harness against each quantization level on representative production traffic — not trusting a vendor’s aggregate benchmark number, which can hide task-specific regressions a general-purpose eval never surfaces.
Failure Modes
Section titled “Failure Modes”KV cache exhaustion under bursty concurrent load
If admission control doesn’t account for the actual remaining KV cache budget (this module’s Implementation section), a burst of concurrent long-context requests can exceed available GPU memory, forcing the server to preempt or reject in-flight requests — a cascading failure that looks like the model server crashed, when it was actually a capacity-planning gap.
Head-of-line blocking from static batching
Under static batching, one unusually long generation in a batch holds every other request in that batch hostage until it finishes, even though their own generations completed long ago. This shows up as unpredictable p99 latency that has nothing to do with an individual request’s own size — the fix is continuous batching (this module’s Architecture section), not a bigger GPU.
Cold-start latency from GPU autoscaling
Scaling a GPU-backed model server from zero (or scaling out under load) means loading multi- gigabyte model weights onto a freshly provisioned GPU before it can serve a single request — this can take tens of seconds to minutes, an eternity compared to the latency budget of the request that triggered the scale-out. See this module’s Scaling section for mitigations.
Quantization accuracy loss going unmeasured
Shipping a lower-precision quantization level purely because it fits more requests per GPU, without re-running the evaluation harness against it, risks a silent quality regression on exactly the tasks a general benchmark doesn’t cover — the same evaluation discipline Module 4 requires for any model change applies here too.
Trade-offs
Section titled “Trade-offs”Tensor parallelism vs. pipeline parallelism
Tensor parallelism lowers per-token latency but demands fast GPU interconnect and pays a synchronization cost on every layer. Pipeline parallelism tolerates weaker interconnect but introduces pipeline bubbles that need a large enough batch to hide.
Quantization level: accuracy vs. concurrent capacity
Lower precision (INT8, INT4) increases how many concurrent requests fit in the same GPU memory and speeds up the memory-bandwidth-bound decode phase, at the cost of some accuracy — real only when measured against an evaluation harness on representative tasks, not assumed from a vendor benchmark.
Keep-warm capacity vs. cost during GPU autoscaling
Keeping GPUs warm and idle avoids cold-start latency entirely but pays for capacity that may go unused; scaling to zero saves cost but reintroduces the cold-start problem on the next request after an idle period. The right point depends on how spiky real traffic actually is, not a default.
Security
Section titled “Security”- Multi-tenant serving must isolate KV cache and batch composition between tenants — a shared serving fleet with no isolation boundary risks one tenant’s traffic pattern (a burst of huge prompts) starving another tenant’s latency budget, or, in a poorly isolated implementation, leaking timing information about other tenants’ request shapes.
- Quantized or distilled models deployed without re-running safety evaluations can regress on safety-relevant behavior that a pure accuracy benchmark wouldn’t catch — treat a quantization change as a model change requiring the same safety review as swapping model versions entirely.
Performance
Section titled “Performance”- Decode throughput, not prefill throughput, usually dominates the cost-per-token calculation for typical chat-style workloads, because output tokens usually outnumber input tokens processed per decode step — this is why continuous batching (which specifically improves decode-phase GPU utilization) has an outsized effect on serving cost.
- Time-to-first-token is bounded by prefill latency; time-per-output-token is bounded by decode latency — these are different metrics with different bottlenecks, and conflating them into a single “latency” number hides which phase actually needs optimizing. This is the same layered- latency discipline Module 3 applies to network requests, applied here to the two phases inside model serving.
- KV cache memory, not raw compute, is usually the binding constraint on how many concurrent requests a GPU can serve — PagedAttention-style memory management (this module’s Deep Dive) exists specifically to relax this constraint.
Scaling
Section titled “Scaling”- Autoscaling GPU-backed fleets must account for cold-start latency, not just request volume — scaling reactively on current load guarantees the new replica isn’t ready until well after the load spike that triggered it. Practical mitigations: predictive/scheduled scale-out ahead of known traffic patterns, and keeping a small warm pool sized to absorb burst latency while cold replicas spin up.
- Horizontal scaling (more replicas behind a load balancer) is the default lever once a single replica is saturated — the load balancer in this module’s Architecture diagram needs its own routing awareness of each replica’s current KV cache utilization, not just round-robin, or it can route a large request to an already-near-capacity replica.
- Model sharding via tensor or pipeline parallelism is a separate scaling axis from replica count — it’s what to reach for when a single model doesn’t fit on one GPU at all, not a substitute for adding replicas once it does fit.
Interview Questions
Section titled “Interview Questions”Why is LLM decode memory-bandwidth-bound instead of compute-bound?
Each decode step generates one token per sequence, and doing so requires reading the entire model’s weights and that sequence’s KV cache from GPU memory — a large memory transfer for a tiny amount of actual computation. Prefill has the opposite ratio, so batching decode steps across many requests is how serving systems recover compute utilization decode alone can’t achieve.
How does continuous batching improve on static batching?
Static batching holds a fixed set of requests until the entire batch finishes, wasting GPU capacity on any slot whose request finished early. Continuous batching re-evaluates admission at every decode step, immediately replacing a finished request’s slot with a queued one, keeping GPU utilization high independent of how uneven request lengths are.
What actually caps concurrent request capacity on a single GPU?
In most production serving setups it’s available KV cache memory, not raw compute — every concurrent sequence needs its own growing KV cache, and that competes with model weights for the same finite GPU memory. PagedAttention-style paged allocation (this module’s Deep Dive) directly targets this constraint by eliminating fragmentation waste in how that memory is allocated.
When would you choose pipeline parallelism over tensor parallelism?
When GPU interconnect bandwidth is limited — tensor parallelism needs fast interconnect because every layer requires cross-GPU synchronization, while pipeline parallelism only needs to pass activations between adjacent stages, tolerating weaker interconnect at the cost of pipeline bubbles that need sufficiently large batches to hide.
How would you mitigate cold-start latency in a GPU-autoscaled model server?
Keep a small warm pool of replicas sized against typical burst shape rather than scaling from zero reactively, and/or scale ahead of known traffic patterns predictively — reactive scale-out on current load structurally guarantees the new replica arrives after the spike that triggered it, per this module’s Failure Modes section.
Hands-on Lab
Section titled “Hands-on Lab”Dynamic Batching Inference implements this module’s
batching mechanism as running code: a batch that closes on whichever of size or timeout fires first,
canary version routing with independent per-version metrics, promote/rollback as routing changes,
and a p50/p95/p99 benchmark harness. It is labelled production-shaped — the mechanism is real, but
the model call is a fixed-cost sleep, which the lab is explicit about hiding the padding-waste curve
that makes optimal batch size a property of the workload.
Simulate head-of-line blocking under static batching
Extend this module’s ContinuousBatchScheduler with a StaticBatchScheduler that only admits new
requests once every in-flight request has finished, run both against the same mixed workload of
short and long max_new_tokens requests, and measure completion time for the short requests under
each scheduler — quantify the head-of-line blocking this module’s Failure Modes section describes,
rather than just asserting it exists.
References
Section titled “References”- Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention” — the vLLM paper.
- Yu et al., “Orca: A Distributed Serving System for Transformer-Based Generative Models” — introduced iteration-level (continuous) batching.
- Module 3: Networking — the layered-latency-budget discipline this module applies to prefill vs. decode.
- Module 4: AI Infrastructure — the evaluation harness this module’s Production Example and Security section apply to quantization decisions.
Revision History
Section titled “Revision History”| Version | Date | Change |
|---|---|---|
| 1.0.0 | 2026-08-07 | Initial publication. |