Module 10: Kubernetes
Read the transcript
1. Why Kubernetes Is Correctness-Critical for AI
Host: So we’re spending this whole episode on Kubernetes and AI, and I want to start by pushing back on the obvious take, which is that this is just generic container orchestration with some GPUs sprinkled in. Why does it deserve that much airtime?
Guest: Because Kubernetes isn’t just running containers here, it’s making the actual decisions that determine whether your AI platform works at all. It decides which physical GPU a workload lands on, how many replicas of a model server exist right now, and whether a rollout can happen without dropping requests mid-flight. Those are expensive, consequential decisions, and Kubernetes will execute them exactly as declared, no matter how wrong that declaration is.
Host: And that’s the scary part, right, it doesn’t fail loudly when you get it wrong. It just quietly does the wrong thing.
Guest: Exactly, it’ll happily schedule a GPU workload onto a node that can never satisfy it and leave it queued forever, or roll a deployment straight through every healthy replica because nothing told it not to. None of this requires exotic knowledge, it’s the same handful of primitives — scheduling, probes, disruption budgets — just applied correctly or not.
2. The Reconciliation Model: Pods, Deployments, Services, HPA, PDB
Host: So let’s lay out the actual pieces Kubernetes is reconciling, because I think people hear ‘orchestration’ and imagine something more active than it is. Walk me through the objects.
Guest: It’s a small composable set. A Pod is just the running container, the thing the scheduler places on a node. A Deployment declares ‘N copies of this pod should exist,’ it creates a ReplicaSet, and that ReplicaSet is the actual loop constantly checking that the count matches — the Deployment layer on top is what orchestrates rolling updates across ReplicaSets. Then a Service gives you a stable network identity in front of whichever pods happen to exist right now, an HPA adjusts that replica count based on metrics, and a PodDisruptionBudget constrains what the cluster is allowed to do to you voluntarily — rollouts, drains — separate from a node just crashing on its own.
Host: And that reconciliation loop is exactly where the GPU danger lives, right — it’s matching desired state to actual state, but it has no idea ‘desired state’ should mean a working GPU.
Guest: Right, it only knows what’s in the manifest. Skip the GPU resource request and a pod can land on a GPU node and just never touch the GPU, silently wasting the most expensive thing in your cluster. Get the request right but forget the toleration for a tainted GPU pool, and now it sits Pending forever, because nothing told it that pool is eligible. Same reconciliation loop either way — it’s doing exactly what you wrote, not what your workload needed.
3. Two Independent Loops: Scheduler vs. HPA
Host: So there are actually two separate control loops running here, and they don’t really talk to each other directly. Walk me through how the scheduler and the HPA end up in the same story.
Guest: Right, the scheduler only runs once, at pod creation — it binds a pod to a node and then it’s done, it’s not watching that decision over time. The HPA is completely different, it’s continuously watching live metrics and adjusting the Deployment’s desired replica count on an ongoing basis. They only interact indirectly: the HPA decides you need more replicas, creates that intent, and then hands the scheduler a fresh batch of pods it now has to go find homes for.
Host: And that’s the part that sounds fine on a dashboard — replica count goes up, looks like a successful scale-up.
Guest: Exactly, but if the GPU pool is already maxed out, those new pods just sit there Pending, they never get bound to a node. The Deployment will happily report that replica count went up, HPA did its job, but your actual serving capacity hasn’t moved at all — you’ve got a metric that looks like success and a user-facing request queue that’s still backing up.
4. GPU Scheduling Correctness: Requests, Limits, Taints, and Tolerations
Host: So that Pending pod problem — is that just capacity, or is there a scheduling correctness issue underneath it too? Because I’ve seen GPU pods land on the wrong node entirely, or refuse to schedule on a node that clearly has a GPU sitting idle.
Guest: That’s almost always requests and limits and the taint setup being wrong, not capacity. The scheduler only looks at what a pod requests to decide if it fits — so if you request nvidia.com/gpu as ‘1’, you get a whole GPU or nothing, there’s no fractional GPU scheduling like there is with CPU. If you get that request wrong, or leave it off entirely, the scheduler doesn’t know it needs to reserve a GPU at all and can over-pack the node, so pods end up competing for actual GPU resources at runtime even though the scheduler thought there was room.
Host: And that’s where the taints come in, right — keeping things from accidentally landing on GPU nodes in the first place?
Guest: Exactly, clusters taint the GPU pool, something like nvidia dot com slash gpu equals present, no schedule, so nothing lands there by default because those nodes are expensive. Your model-server pod then needs a toleration for that taint plus, usually, a nodeSelector or affinity rule pointing at the GPU pool by label — the toleration just says ‘I’m allowed here,’ it doesn’t say ‘put me here,’ so you want both declared explicitly rather than trusting one mechanism to do the whole job.
5. Readiness vs. Liveness: Don’t Kill What You Should Just Pause
Host: So we’ve got the toleration and affinity sorted, the pod’s actually landing on the right GPU node. But there’s a whole separate failure mode once it’s running: what happens when it’s temporarily struggling versus actually dead. Walk me through readiness versus liveness, because I’ve seen these two get conflated in ways that are genuinely expensive here.
Guest: Right, they answer completely different questions. Liveness is ‘should this process be killed and restarted,’ readiness is ‘should this pod currently get traffic’ — and failing readiness just pulls it out of the Service’s endpoint list without touching the process at all. The failure mode is a pod under a heavy batch of inference requests failing a liveness check because it’s slow to respond, so Kubernetes kills it and you lose every in-flight request, when a readiness probe would’ve just paused new traffic for a few seconds until it caught up.
Host: And the example in the module doesn’t hardcode /readyz to always return 200, it’s actually wired to real state — there’s a draining flag that flips on shutdown.
Guest: Exactly, that’s the part people skip. On shutdown you set draining to true and readyz immediately starts returning 503, so Kubernetes stops routing new requests before SIGTERM even finishes tearing the pod down — you get a truthful signal instead of a lie. And that’s literally the same graceful-drain discipline as the DrainCoordinator pattern from the gateway module; the readiness probe is just that same idea surfaced at the layer Kubernetes actually checks.
6. Rolling Updates and PodDisruptionBudgets
Host: Okay, so readiness handles the single-pod case truthfully. But what stops a rollout or a node drain from just yanking multiple replicas at once, even if each one is individually reporting its state correctly?
Guest: That’s what maxUnavailable and maxSurge on the Deployment, and the PodDisruptionBudget, are for. The rollout config controls how many old pods get replaced at once during your own deploy, but a PDB with minAvailable set to two, say, is a floor that applies to any voluntary disruption — your rollout, but also a node drain for maintenance or an autoscaler cycling nodes out. Kubernetes literally refuses to evict a pod if doing so would drop you below that floor.
Host: So without the PDB, a node drain during a deploy could take down every replica at the same time if they happened to land on the same node?
Guest: Right, that’s exactly the scenario it prevents — three replicas, no PDB, all co-located, and a drain event or a badly timed rollout takes all three down simultaneously because nothing told Kubernetes that’s unacceptable. The PDB is the thing that makes ‘at least two of these must stay up’ an actual constraint the eviction and rollout machinery has to respect, not just an assumption you’re hoping holds.
7. StatefulSets: When Replicas Aren’t Interchangeable
Host: So PDBs handle the ‘don’t take them all down at once’ problem, but that still assumes any surviving replica can pick up the slack. What happens when that assumption itself is false — when replica three isn’t actually a substitute for replica one?
Guest: Then you’ve got the wrong object entirely. A Deployment gives you interchangeable pods — any replica can be killed and rescheduled and a new one comes up with no identity, no history, and that’s fine because any replica can serve any request. But a sharded vector index or a model server with pinned GPU-to-shard assignment doesn’t work that way — pod-2 owns shard 2’s data and state, and if it gets rescheduled with no memory of being pod-2, it doesn’t resume, it just shows up empty. That’s what a StatefulSet is for: stable ordinal identity, pod-0 through pod-n, and storage that follows each specific pod across rescheduling instead of being handed out fresh.
Host: And the failure mode runs both directions, presumably — it’s not just ‘stateful workload as a Deployment loses your shards.’
Guest: Exactly, the reverse mistake is just as real: taking a genuinely stateless model server and running it as a StatefulSet out of caution buys you nothing but the ordinal identity and per-pod storage overhead, with no benefit since there’s no state to reattach. So this isn’t ‘StatefulSets are safer, default to them’ — it’s match the object to whether replicas are actually interchangeable. Ask does pod-2 need to come back as pod-2, with its data, or can literally any fresh pod take its place — that answer tells you which object you need, and getting it backwards either leaves a rescheduled pod unable to resume correctly because it lost its stable identity and storage, or just wastes operational complexity you didn’t need to take on.
8. The Illusion of Successful Autoscaling
Host: Let’s walk through a scenario that I think trips up a lot of teams: the HPA does exactly what it’s supposed to do, scales from 3 to 20 replicas under load, and the dashboard proudly shows 20. What’s actually going wrong here?
Guest: The HPA created 20 pod specs, but GPU capacity is scarce and slow to provision, so the cluster autoscaler either hasn’t spun up new GPU nodes yet or physically can’t because of quota limits or a regional shortage. Those new pods just sit there in a Pending state indefinitely, waiting for a node that isn’t coming any time soon. Meanwhile the HPA’s own reporting only tracks desired replica count against its metric, so from its point of view it did its job correctly — the number says 20, and everything looks fine unless you know to look past that number.
Host: So the dashboard is technically accurate and completely misleading at the same time. What actually catches this before it becomes an incident?
Guest: You have to check pod status directly — Pending versus Running — not just trust the HPA’s reported replica count, because that count reflects intent, not served capacity. And it compounds with two other things worth checking: is the HPA even scaling on the right metric, since CPU utilization on a GPU-bound serving workload can lag real demand, and is the node pool’s autoscaler ceiling actually large enough to satisfy maxReplicas in the first place. GPU provisioning commonly takes minutes rather than the seconds CPU capacity takes, which is exactly why keep-warm buffers matter so much more here than for a typical stateless web service.
9. Security as Isolation: Resource Limits, Taints, and RBAC
Host: Let’s shift gears slightly — we’ve talked about requests and limits and taints purely as scheduling mechanics, getting the right pod on the right node. But you’ve said there’s a security dimension to all of this that people miss. What do you mean?
Guest: Those same primitives are your isolation boundary, whether you designed them that way or not. If you don’t enforce resource limits on a multi-tenant cluster, one noisy or runaway workload degrades everything else sharing that node — that’s not a scheduling bug, that’s a tenant-isolation failure. Same logic applies to GPU taints: without them, any pod with a loose enough scheduling policy can land on expensive GPU capacity it was never authorized to touch, so the taint is functioning as coarse access control, not just placement hygiene. And it extends to identity too — if you’ve got an operator or controller pod calling the Kubernetes API directly, its service account needs least-privilege RBAC scoping just like any other credential, because the default service account carrying cluster-admin-equivalent permissions is a much bigger blast radius than people account for when they’re just trying to get something scheduled.
10. Trade-offs: GPU Pool Topology and the Scale-Down Gamble
Host: Let’s talk cluster topology for a second, because I think there’s a real dedicated-versus-mixed-pool decision buried in here. Why not just let GPU and non-GPU workloads share nodes and let the scheduler sort it out?
Guest: You can, but you’re trading cost attribution and safety for simplicity. A dedicated tainted GPU pool means only GPU workloads land there, your billing is clean, and nothing accidentally eats scheduling headroom on your expensive nodes — the cost is you’re now managing an extra pool and you can end up with idle GPU capacity if your bin-packing isn’t tight. A mixed pool is one less thing to operate, but a CPU-heavy workload’s resource requests can crowd out the room a GPU pod needs, and now you’re debugging a scheduling failure that’s really a topology decision you made months ago. There’s a similar gamble on the scaling side: scale down fast after a load drop and you save real money, but you’re betting the next traffic spike doesn’t arrive before you’ve re-provisioned GPU capacity, and that provisioning delay is minutes, not seconds — so you repay the cold-start penalty you thought you’d already gotten past.
Host: So keeping some replica headroom at rest is basically buying insurance against your own autoscaler being too good at its job.
Guest: Exactly, and the same speed-versus-capacity trade shows up in rollouts through maxSurge — crank it up and your rolling update finishes faster, but you need more total capacity sitting available during the transition, which on GPU nodes is not a cheap ask. None of these are right-or-wrong settings, they’re dials, and the mistake is leaving them at a stateless-web-service default and assuming GPU economics don’t apply.
11. Proving It Yourself: The Lab Reproduction and the Interview Questions
Host: So before we let people go, there’s an exercise worth actually doing rather than just nodding along to. Take this module’s Deployment manifest, unmodified, and apply it to a local kind or minikube cluster that has no real GPU nodes at all. Then just check on your pods and watch what happens.
Guest: Every replica sits Pending, and that’s the whole point — the node selector field and the tolerations are targeting a GPU pool that doesn’t exist on your laptop, so the scheduler has nowhere to place them and it just says so, quietly, forever. Then go back in, strip out the node selector, the tolerations, and the GPU resource request, reapply, and watch the same pods schedule immediately. You’ve now reproduced, on purpose, in five minutes, the exact failure this whole episode has been warning about — and if you ever get asked in an interview what you check when an HPA says it wants five replicas but only two are serving traffic, the answer is exactly this: check pod status, not the HPA’s desired count, because Pending means the scheduler is stuck, not that the app is broken.
Host: Which is really the thread running through every segment today — liveness versus readiness, PDBs during a drain, StatefulSets for pinned shards, the scale-down gamble on GPU pools — none of it is Kubernetes being clever on your behalf. It does precisely what the manifest says, so the job is making sure the manifest says what you actually mean. That’s the episode — thanks for listening, and go make those pods sit Pending on purpose.
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Executive Summary
Section titled “Executive Summary”Kubernetes matters to an AI platform for a specific reason beyond general container orchestration: it’s the layer that decides which physical GPU a workload lands on, how many replicas of a model server exist right now, and whether a rollout can proceed without dropping in-flight requests. Get the scheduling, probe, and disruption-budget mechanics wrong and Kubernetes will confidently do the wrong thing — schedule a GPU workload onto a CPU node’s request queue forever, or roll a deployment straight through every replica capable of serving traffic. None of this is exotic; it’s the same handful of primitives, applied correctly.
Mental Model
Section titled “Mental Model”Kubernetes doesn’t run anything directly — it continuously reconciles observed cluster state toward declared desired state, through a small set of objects that compose:
- Pod — the actual running container(s); the unit the scheduler places onto a node.
- Deployment → ReplicaSet → Pod — declares “N copies of this pod should exist”; the ReplicaSet is the reconciliation loop that keeps the pod count correct, the Deployment is what manages rolling updates across ReplicaSets.
- Service — a stable network identity in front of a changing set of pods.
- HorizontalPodAutoscaler (HPA) — adjusts a Deployment’s replica count based on observed metrics.
- PodDisruptionBudget (PDB) — a constraint the cluster must respect during voluntary disruptions (rollouts, node drains), separate from involuntary ones (a node crashing).
The scheduler only knows what you declare, not what you mean
A pod without a GPU resource request can be scheduled onto a GPU node and silently never use the
GPU; a pod with a GPU request but no matching toleration can sit Pending forever if the GPU pool
is tainted. Kubernetes does exactly what the manifest says, not what the workload actually needs
— which is why declaring resource requests and tolerations correctly is a correctness requirement
for AI workloads, not an optimization to add later.
Architecture
Section titled “Architecture”flowchart TB
Client[Client request] --> Ingress[Ingress]
Ingress --> Svc[Service]
Svc --> Pod1[Pod]
subgraph Workloads["Workload objects"]
HPA["HorizontalPodAutoscaler (watches CPU / GPU / custom metrics)"] --> Deploy[Deployment]
Deploy --> RS[ReplicaSet]
RS --> Pod1
PDB["PodDisruptionBudget (min available)"] -.protects during drain.-> Pod1
end
subgraph ControlPlane["Control plane"]
API[API server] --> Etcd[(etcd)]
Sched[Scheduler] --> API
end
RS -->|desired replica count| API
subgraph Nodes["Node pools"]
subgraph CPUPool["CPU node pool"]
NodeCPU[Node]
end
subgraph GPUPool["GPU node pool: tainted + nodeSelector"]
NodeGPU[Node]
end
end
Sched -->|binds pod, no GPU request| NodeCPU
Sched -->|binds pod, requests nvidia.com/gpu + toleration| NodeGPU
Pod1 -.scheduled onto.-> NodeCPUTwo independent reconciliation loops are visible in this diagram: the scheduler binding pods to
nodes (once, at pod creation, based on resource requests and node taints/tolerations) and the
HPA adjusting the Deployment’s desired replica count on an ongoing basis (based on live metrics).
They interact indirectly — a scale-up event from the HPA creates new pods that the scheduler then
has to place, and if the GPU pool has no capacity left, those new pods stay Pending instead of
serving traffic, an autoscaling event that looks successful (replica count went up) but delivers no
actual capacity increase.
Deep Dive
Section titled “Deep Dive”Resource requests and limits. A request is what the scheduler uses to decide whether a pod
fits on a node (and reserves that resource on the node accordingly); a limit is the hard ceiling
the container is throttled or killed at. Setting a request too low lets the scheduler over-pack a
node, causing pods to compete for actual resources at runtime even though the scheduler thought
there was room; setting a limit lower than the request is invalid configuration, not a safety net.
For GPU workloads, nvidia.com/gpu is typically requested as an integer (fractional GPU requests
aren’t supported the way CPU/memory are) — a pod either gets a whole GPU or none.
Taints and tolerations, node affinity. GPU nodes are expensive, so clusters typically taint the
GPU node pool (nvidia.com/gpu=present:NoSchedule) so nothing lands there by accident; a pod that
actually needs a GPU declares both the resource request and a matching toleration, and usually a
nodeSelector or node affinity rule to be explicit rather than relying on the taint/toleration
mechanism alone to route it correctly.
Readiness vs. liveness probes. A liveness probe answers “is this process alive, or should it be restarted?” — failing it kills and restarts the container. A readiness probe answers “should this pod currently receive traffic?” — failing it removes the pod from the Service’s endpoint list without restarting it. Conflating the two is a common failure: a pod under temporary load that fails a liveness probe gets restarted (losing all in-flight requests) when a readiness probe would have correctly just paused new traffic to it until it recovered.
Rolling updates and PodDisruptionBudgets. A Deployment’s rolling update replaces old pods with
new ones gradually, controlled by maxUnavailable/maxSurge. A PodDisruptionBudget sets a floor
(minAvailable) or ceiling (maxUnavailable) on how many pods can be down at once for voluntary
disruptions — both a rollout and a node drain (for cluster maintenance or autoscaler node
replacement) respect it, refusing to proceed if it would violate the budget. Without a PDB, a node
drain during a deploy can take down every replica of a service simultaneously if they happen to be
co-located.
StatefulSets. Deployments treat pods as interchangeable; a StatefulSet gives each pod a stable,
ordinal identity (pod-0, pod-1, …) and stable storage that follows it across rescheduling.
This matters for anything that isn’t horizontally stateless by default — a self-hosted vector index
or a model-serving setup relying on sharded, pinned GPU assignment per replica, where “any replica
can handle any request” (the Deployment assumption) doesn’t hold.
Research Note
Primary references for scheduling constraints and disruption budgets — worth reading directly because the exact interaction between taints, tolerations, and node affinity is easy to misremember, and it’s exactly the kind of detail interview questions target.
Source: Kubernetes documentation, "Assigning Pods to Nodes" and "Pod Disruption Budgets"
Implementation
Section titled “Implementation”The Deployment, HPA, and PodDisruptionBudget for a GPU-backed model server, showing the
resource-request/toleration pairing this module’s Deep Dive covers, followed by a readiness probe
implementation wired to real in-process drain state rather than a hardcoded 200 OK:
GPU-scheduled Deployment with autoscaling and a disruption budget
apiVersion: apps/v1kind: Deploymentmetadata: name: model-serverspec: replicas: 3 selector: matchLabels: { app: model-server } template: metadata: labels: { app: model-server } spec: tolerations: - key: nvidia.com/gpu operator: Exists effect: NoSchedule nodeSelector: cloud.provider/pool: gpu containers: - name: model-server image: registry.example.com/model-server:1.4.0 resources: requests: { cpu: "2", memory: "8Gi", nvidia.com/gpu: "1" } limits: { cpu: "4", memory: "16Gi", nvidia.com/gpu: "1" } readinessProbe: httpGet: { path: /readyz, port: 8080 } periodSeconds: 5 livenessProbe: httpGet: { path: /healthz, port: 8080 } periodSeconds: 15---apiVersion: autoscaling/v2kind: HorizontalPodAutoscalermetadata: name: model-serverspec: scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: model-server } minReplicas: 3 maxReplicas: 20 metrics: - type: Pods pods: metric: { name: inflight_requests_per_pod } target: { type: AverageValue, averageValue: "8" }---apiVersion: policy/v1kind: PodDisruptionBudgetmetadata: name: model-serverspec: minAvailable: 2 selector: matchLabels: { app: model-server }The /readyz endpoint this Deployment probes should reflect real application state, not always
return success:
from __future__ import annotations
from fastapi import FastAPI, Response
app = FastAPI()draining = False
@app.get("/readyz")async def readyz() -> Response: if draining: return Response(status_code=503, content="draining") return Response(status_code=200, content="ready")
@app.on_event("shutdown")async def on_shutdown() -> None: global draining draining = TrueSetting draining = True on shutdown and failing /readyz immediately gives Kubernetes a truthful
signal to stop routing new traffic to this pod before SIGTERM finishes tearing it down — the same
graceful-drain discipline Module 1’s
DrainCoordinator implements at the application level, now surfaced through the readiness probe
Kubernetes actually checks.
Production Example
Section titled “Production Example”A model-serving Deployment scales from 3 to 20 replicas under load, per the HPA manifest above. If
the GPU node pool’s autoscaler can’t provision new GPU nodes fast enough (GPU capacity is often
scarcer and slower to provision than CPU capacity), the new pods the HPA created sit Pending
indefinitely — the HPA reports the “correct” desired replica count, dashboards show 20 replicas
configured, and yet actual serving capacity hasn’t increased at all. Diagnosing this requires
checking pod status (Pending vs Running), not just the HPA’s reported replica count, which is
precisely why this module’s Failure Modes section treats this as its own distinct failure rather
than folding it into generic “autoscaling.”
Failure Modes
Section titled “Failure Modes”GPU pods stuck Pending after a scale-up
The HPA increases desired replicas, new pods are created, but the GPU node pool has no available capacity and the cluster autoscaler hasn’t provisioned a new GPU node yet (or can’t — GPU quota limits, regional GPU shortages). The Deployment’s replica count looks correct while actual serving capacity hasn’t moved; this module’s Production Example walks through exactly this scenario.
Liveness probe misconfigured as a readiness check
A liveness probe that fails under temporary load (a slow but recovering model server) triggers a restart, discarding every in-flight request that pod was serving — when a readiness probe failing would have simply paused new traffic without destroying existing work. This module’s Deep Dive section on probes is the fix: know which failure mode each probe type produces before wiring either one to an endpoint that can legitimately be slow sometimes.
Rolling update without a PodDisruptionBudget takes down every replica at once
If a node drain (cluster maintenance, autoscaler node replacement) coincides with a rollout and no
PDB constrains it, and multiple replicas of the same service happen to be co-located on the
draining node, all of them can go down simultaneously — a full outage that a minAvailable PDB
would have prevented by making the drain wait.
Resource requests set too low for the workload's real usage
A pod whose CPU/memory request understates what it actually needs lets the scheduler pack more pods per node than the node can actually sustain under real load — the failure shows up as throttling or OOM-kills under production traffic that never appeared in a lightly loaded staging environment, because staging never triggered the over-packing.
StatefulSet assumptions applied to a workload that doesn't need them, or vice versa
Running a genuinely stateful workload (a self-hosted vector index shard, per Module 8) as a Deployment loses the stable identity and storage a rescheduled pod needs to resume correctly; running a genuinely stateless model server as a StatefulSet adds ordinal-identity and storage overhead for no benefit. Match the object to the workload’s actual state requirements, not habit.
Trade-offs
Section titled “Trade-offs”Separate GPU node pool vs. mixed node pool
A dedicated, tainted GPU node pool keeps GPU capacity from being accidentally consumed by non-GPU workloads and makes GPU cost attribution clean, at the cost of an extra pool to manage and potential GPU idle time if scheduling isn’t tight. A mixed pool simplifies cluster topology but risks a non-GPU workload’s resource requests crowding out GPU-workload scheduling headroom.
Aggressive HPA scale-down vs. keeping headroom
Scaling down quickly after load drops saves cost but risks scaling back up into the same GPU provisioning delay this module’s Failure Modes section describes, on the very next traffic spike. Keeping some replica headroom costs more at rest but avoids repeatedly re-paying the GPU cold-start penalty Module 9 already covers.
Deployment vs. StatefulSet for model servers
A Deployment is simpler to operate and scale when every replica is truly interchangeable. A StatefulSet is required the moment replicas aren’t interchangeable — sharded state, pinned GPU-to-shard assignment — at the cost of the extra operational complexity ordinal identity and per-pod storage bring.
Security
Section titled “Security”- Resource requests/limits are a tenant-isolation control, not just a scheduling hint — a multi-tenant cluster without enforced limits lets one workload’s runaway resource usage degrade every other workload scheduled on the same node.
- Node taints for GPU pools double as a coarse access-control boundary — without them, any pod with a sufficiently permissive scheduling policy could land on expensive GPU capacity it was never authorized to consume.
- Service accounts and RBAC scoping for pods that call the Kubernetes API directly (an operator, a controller) need the same least-privilege discipline as any other credential — a pod’s default service account should not carry cluster-admin-equivalent permissions by default.
Performance
Section titled “Performance”- Probe intervals and timeouts directly affect failover speed — a readiness probe with a long
periodSecondsdelays how quickly Kubernetes notices a pod should stop receiving traffic, extending the window of failed requests during a real problem. - GPU node provisioning latency is usually the dominant term in autoscaling response time for AI workloads specifically — CPU autoscaling can provision new capacity in seconds; GPU capacity commonly takes minutes, which is why keep-warm strategies matter more here than for typical stateless web services (see Module 9).
- Rolling update
maxSurge/maxUnavailablesettings trade rollout speed for spare capacity — a largermaxSurgefinishes a rollout faster but temporarily needs more total capacity available.
Scaling
Section titled “Scaling”- The HPA’s target metric matters more than its existence — scaling on CPU utilization for a request-serving workload whose actual bottleneck is in-flight-request count or GPU memory (per Module 9) scales on the wrong signal and lags real demand.
- Cluster autoscaler node-pool limits are a hard ceiling independent of the HPA’s
maxReplicas— an HPA configured for 20 replicas is meaningless if the GPU node pool’s autoscaler is capped at capacity for 10; both limits need to be sized against the same capacity plan. - Multi-region or multi-cluster scaling moves the scaling question up a level — from “how many pods in this cluster” to “how much traffic does each region’s cluster serve,” a decision this module’s Architecture stays within a single cluster’s boundary and Module 11: Cloud picks up.
Interview Questions
Section titled “Interview Questions”A Deployment's HPA reports the desired replica count increased, but latency hasn't improved. How do you debug it?
Check actual pod status, not just the HPA’s desired-replica number — Pending pods mean the
scheduler can’t place them, most commonly because the node pool (especially a GPU pool) has no
available capacity and the cluster autoscaler hasn’t (or can’t) provision more nodes. This module’s
Production Example walks through exactly this failure.
What's the difference between a readiness probe and a liveness probe, and why does mixing them up matter?
A liveness probe failure restarts the container, discarding in-flight work; a readiness probe failure only pauses new traffic to the pod. Wiring a liveness probe to an endpoint that can legitimately be slow under normal load causes unnecessary restarts that destroy in-flight requests a readiness probe would have handled safely.
How do PodDisruptionBudgets interact with node drains and rolling updates?
A PDB sets a floor or ceiling on how many pods of a selector can be down simultaneously for voluntary disruptions. Both a rolling update and a node drain check the PDB before proceeding and will pause rather than violate it — without one, a node drain that happens to hold multiple replicas of the same service can take all of them down at once.
When would you use a StatefulSet instead of a Deployment for a model-serving workload?
When replicas aren’t interchangeable — a sharded, self-hosted vector index or any setup with pinned GPU-to-shard assignment needs the stable ordinal identity and storage a StatefulSet provides. A stateless model server where any replica can handle any request has no need for it and should stay a Deployment.
Hands-on Lab
Section titled “Hands-on Lab”No dedicated Kubernetes lab exists yet — see the Roadmap for what’s planned.
Reproduce GPU pods stuck Pending
On a local cluster with no real GPU nodes (kind or minikube), apply this module’s Deployment
manifest unmodified — its nodeSelector and tolerations target a GPU pool that doesn’t exist —
and observe every replica sit Pending via kubectl get pods. Then remove the GPU-specific
nodeSelector, tolerations, and nvidia.com/gpu resource request and reapply, confirming the
pods schedule successfully — reproducing, on purpose, the exact mismatch this module’s Mental
Model section warns about.
References
Section titled “References”- Kubernetes documentation, “Assigning Pods to Nodes”
- Kubernetes documentation, “Pod Disruption Budgets”
- Kubernetes documentation, “Configure Liveness, Readiness and Startup Probes”
- Module 1: Production Python — the
DrainCoordinatorapplication-level drain logic this module’s readiness probe surfaces to Kubernetes. - Module 9: Model Serving — the GPU cold-start problem this module’s autoscaling sections build on.
Revision History
Section titled “Revision History”| Version | Date | Change |
|---|---|---|
| 1.0.0 | 2026-08-07 | Initial publication. |