Architecture: Multi-Region AI Serving Failover
Read the transcript
1. The diagram stops where the work starts
Host: So picture the diagram everybody draws for AI serving failover: box one is your primary region, box two is the standby, there’s an arrow between them, and somebody writes an RTO next to it like fifteen minutes. Clean, satisfying, done. Today we’re arguing that diagram is where the real work hasn’t even started yet.
Guest: Right, because that RTO isn’t a fact about your system, it’s a budget you haven’t spent. Something has to detect the region is actually down, something has to promote the standby, traffic has to get redirected, and for AI serving specifically, the standby has to get the model loaded onto GPUs before it can serve a single request. Those happen in sequence, not in parallel, and the diagram just doesn’t show any of that.
Host: So if I write fifteen minutes on a slide, I haven’t actually checked whether fifteen minutes is achievable, I’ve just stated that it’s desirable. We’re going to walk through that sequence stage by stage today, and spend most of our time on the one everyone forgets, which is apparently getting the model warm on the standby’s GPUs.
2. What the trigger has to do, and what makes it hard
Host: Okay, so before we get into warm-up, let’s talk about the trigger itself — the thing that decides ‘yes, failover now.’ What does it actually have to get right?
Guest: Four things, and they pull against each other. It has to commit to that RTO you mentioned, measured from the first failure to the first request served elsewhere. It has to ignore transient faults, because a false failover is its own incident. It has to catch a region that goes fully dark rather than politely reporting itself unhealthy. And it has to behave predictably if the controller making these decisions restarts.
Host: That second one — ignoring transient faults — sounds like it’s directly in tension with speed. How do you tell a dropped packet or a GC pause from the real thing on the first bad probe?
Guest: You can’t, not at the first probe. The only way to tell them apart is to wait and see if it recovers, and every second of that waiting comes straight out of the same RTO budget you’re trying to hit. Combine that with the fact that detection, promotion, traffic shift, and warm-up run one after another, not in parallel, and you can see why the budget gets tight before you’ve even touched the model.
3. Reading the budget left to right
Host: So let’s actually read the thing left to right instead of treating it as one lump number. Detection is the only stage you, the controller, get to pick — everything downstream of that is the platform’s problem, not yours. Is that the right way to think about it?
Guest: Exactly, and that split matters because it tells you what’s negotiable and what isn’t. Everything after detection belongs to the platform, and for AI serving the biggest variable in that fixed cost is the standby tier — hot, warm, or cold. A hot standby is already serving the model, warm has capacity but needs to load weights, cold has to go find capacity before it can even start loading.
Host: So walk me through what that leaves for detection once you actually subtract the platform’s cut from a 15-minute RTO. Give me the lab’s default numbers.
Guest: At their defaults, a hot standby leaves you 720 seconds to detect and decide before you’ve blown the RTO. Warm leaves you 580. Cold goes negative — meaning there’s no detection window at all, no trigger logic you could write, that hits a 15-minute target, because the platform’s own fixed cost already exceeds it.
4. Four ways the trigger gets it wrong
Host: So if the window is the thing eating the budget, the obvious fix is just set the window to the RTO, right? Fifteen minutes, done. Why doesn’t that work?
Guest: Because a window trigger can’t fire until the whole window has been unhealthy — so detection alone consumes the entire budget before promotion even starts. In the lab, a 900-second window against a 900-second RTO overran on every single tier, 1,080 seconds even for a hot standby with nothing to warm. It meets the RTO on paper and misses it by construction.
Host: Okay, that’s one way to get it wrong. What are the other three?
Guest: A controller that treats ‘no checks’ as ‘nothing wrong’ misses a dark region almost every time — in the sweep it caught it in 2 of 30 runs, both by luck. Swing the other way and make it fire on the first bad probe, and it detects instantly but fires 72 times a day on ordinary blips with no outage at all. And a freshly restarted controller has no history, so one unhealthy check looks like the whole window — which fails the region over on its first bad probe, right when restarts and bad probes both cluster: during a deploy.
5. Patience sets both numbers at once
Host: So patience is the knob, and it sets two numbers at once. Walk me through what happens when you turn it.
Guest: In the lab, 3 consecutive failed probes detects in 20 seconds but fires 25 false failovers a day. Push to 12 consecutive and detection slows to 110 seconds, but false failovers drop to 0.2 a day. There’s no setting that’s both fast and quiet — you’re choosing which cost you can absorb, and windows behave almost identically to consecutive counts, since a window of W seconds is just W over interval plus one. Pick whichever shape your team reasons about correctly, because the real design work is the wait time, not the family.
Guest: And none of those numbers travel to your system as-is — triple the mean blip length and that same 12-consecutive rule goes from 0.2 false failovers a day to 13.2. The shape holds, patience buys quiet at the cost of speed, but the magnitudes are only true for the blip distribution they were measured on.
Host: So you can’t borrow someone else’s thresholds, you have to re-fit them against your own probe history before you trust any of these numbers.
6. The standby tier is the first decision, not the last
Host: Okay, so we’ve spent two segments getting the trigger right, but you said the standby tier actually has to be picked first. Why does the order matter?
Guest: Because the tier decision sets the physics the trigger has to live inside. Hot standby means you’re paying for GPU capacity that serves nothing until the day it serves everything. Warm means you’ve got the capacity but not the model loaded. Cold means you’re paying for neither, and you accept the longest warm-up when the day comes.
Host: And that’s not just a one-time cost calculation, you said false failovers have their own hidden price tag.
Guest: Right, and it’s invisible on the bill but very real: every false failover is engineering time scrambling to understand what happened, a cold cache now serving in the region that just took over, and the operational risk that this particular failover goes wrong. A trigger tuned to fire fast on blips is buying those saved seconds with that cost, repeated many times a day, which is why you can’t tune detection until you know what tier you’re protecting.
7. The boundary moves with the traffic
Host: So we’ve covered the tiering and the triggering, but there’s a dimension that doesn’t show up on the budget at all: security. What changes there when you fail over?
Guest: Everything, because the failover carries the security boundary with it. The standby has to hold the same credentials, network policy, and model-artifact access as primary, provisioned through the exact same path — the moment someone hand-adds a role or secret to primary and forgets standby, that’s drift waiting to surface mid-incident. And the controller itself is a juicy target, because anything that can feed it fake unhealthy probes can trigger a failover on command, so that health signal needs the same integrity guarantees as any other control-plane input.
8. Bigger models, more regions, same math
Host: Okay, let’s talk scale, because surely a bigger model or a few more regions just means a bigger number on the same diagram, right? Where does the math actually break?
Guest: The failover decision itself stays cheap, it’s the stages after it that balloon. Warm-up scales with model size, but also with how many replicas are loading weights at once — if your standby is pulling weights for many replicas in parallel, they’re all fighting over the same storage throughput, so the per-replica number you measured in isolation is a lie once you load concurrently. And more regions don’t buy you faster detection at all, they just multiply your homework: every standby target has its own tier, its own warm-up curve, its own contention profile, so a three-region plan isn’t one budget, it’s three budgets, one for each target rather than one for the whole platform.
9. Shipping it: the checklist, what to measure, and what’s deliberately left out
Host: So let’s land this as something a team could actually ship. If you had to write the checklist on one page, what’s on it?
Guest: Standby config gets provisioned through the same path as primary and diffed so it can’t quietly drift — that diffing is what actually catches silent divergence before it matters. Then you exercise the whole thing end to end on a schedule, every stage timed, because a plan you’ve never run is just a diagram with extra steps, and the false-failover rate gets owned by someone who watches it change as traffic does.
Host: And the observability side is basically what makes that checklist enforceable rather than aspirational — you need spans per stage, alerts on missing probes, false-failover rate as a tracked number, warm-up measured live instead of guessed.
Guest: Exactly, and I want to be honest about the edge of this: replication lag at the moment of failure, the RPO side, is a separate problem we’re deliberately not solving here. You can hit your RTO perfectly and still lose more data than you can tolerate, because modeling lag under the load right before an outage basically requires assuming the answer. So measure that separately, don’t let a clean failover drill convince you the data story is fine too — that’s a different budget, with its own math, for another day.
Not covered
The planner wanted these and found nothing in the source to support them:
- Detailed modeling of replication lag / RPO under pre-outage load — the lab explicitly excludes this as it would bake in its own conclusion
- Cost-attribution tagging schemes across teams — covered in Module 11 but not part of the failover design itself
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Problem
Section titled “Problem”A model-serving platform in one region has one region’s availability. The standard answer is a second region and a controller that moves traffic to it when the first one fails. The diagram version of that answer is easy, and it is where most designs stop: two boxes, an arrow, and an RTO written next to it.
The RTO is a budget, and the diagram never spends it. Something has to notice the region is down, something has to promote the standby, traffic has to move, and — the part that is particular to AI serving — the standby has to have the model loaded on GPUs before it can answer anything. Each of those takes time, and they happen in sequence. A design that does not write the sequence out has not checked whether its RTO is achievable, only whether it is desirable.
Module 11 covers the cloud decisions around this. This page covers the failover decision itself, and what the trigger that makes it costs.
Requirements
Section titled “Requirements”- Fail traffic over to a standby region within a committed RTO, measured from the start of the outage to the first request served from the standby.
- Not fail over on transient faults that would have recovered on their own: a failover is an incident in its own right.
- Detect an outage that takes the health signal down with it — a region that goes dark rather than one that reports itself unhealthy.
- Behave predictably when the controller itself restarts.
Constraints
Section titled “Constraints”The stages are sequential. Detection, promotion, traffic shift, and warm-up do not overlap in the general case, so the RTO has to cover their sum. Detection is the only stage whose duration is a policy choice made in the controller; the rest are properties of the platform.
Model warm-up is large and model-shaped. A standby that is not already holding the model in GPU memory has to load it. At the backing lab’s default inputs — 140 GB of weights at 1 GB/s — that is 140 seconds before capacity, and ten minutes more if capacity has to be found first. These are inputs to plug your own numbers into, not measurements; the point is that the term scales with model size, and it is the term that grows when the model does.
Health signals are noisy. Probes fail for reasons that are not region failures: a dropped packet, a GC pause, a deploy. A trigger cannot tell a blip from an outage at the first bad probe — only by waiting — and waiting is drawn from the same RTO budget.
Request Flow
Section titled “Request Flow”flowchart TD
O([Outage begins]) --> D["Detection<br/>measured: the trigger's patience"]
D --> P["Promotion<br/>input"]
P --> T["Traffic shift<br/>input: DNS TTL, reconnects"]
T --> W{"Standby tier"}
W -->|hot| R([Serving again])
W -->|warm| L["Load model weights<br/>input: GB / throughput"]
W -->|cold| C["Find GPU capacity<br/>input"]
C --> L
L --> RRead the diagram as a budget, top to bottom. Detection comes first and is the only stage the controller chooses. Everything after it belongs to the platform, and for AI serving the standby tier decides how much of it there is: a hot standby already serves the model, a warm one has capacity but must load weights, a cold one must find capacity first.
The consequence is the rule this whole page rests on: the detection window is what the RTO leaves after everything else, not the RTO itself. At the backing lab’s default inputs and a 15-minute RTO, that is 720 seconds for a hot standby and 580 for a warm one — and below zero for a cold one, where no trigger at all can meet the target.
Failure Modes
Section titled “Failure Modes”A detection window equal to the RTO
The natural reading of “the trigger should honour the RTO” is to make the window the RTO. A window trigger cannot fire until the whole window has been unhealthy, so detection then consumes the entire budget before promotion has started. In the backing lab, a 900-second window against a 900-second RTO overran on every standby tier — 1,080 seconds even for a hot standby with nothing to warm. The design meets its RTO on paper and misses it by construction.
Silence read as health
A controller that only reasons about the checks it has will treat “no checks” as “nothing wrong”. An outage that takes the prober down with the region produces exactly no checks. In the backing lab, a controller with that behaviour caught a dark region in 2 of 30 runs — both by luck, where the region went dark mid-blip and the last checks it held happened to be unhealthy. Counting a probe that was due and did not arrive as a failure caught all 30.
Failing over on blips
The opposite error. A trigger that fires on the first failed probe detects instantly and fails over constantly: at the lab’s default blip rate it fired 72 times a simulated day with no outage at all. Every one of those moves traffic, cold-starts caches, and pages someone. Eager detection does not remove an incident; it schedules many smaller ones.
A restarted controller with no history
A window trigger with an empty history treats one unhealthy check as every check in the window, so a controller that restarts during a blip fails the region over on its first bad probe. Restarts happen at exactly the moments probes are noisiest — during deploys — which is when this is most likely to fire.
A standby that has drifted
Detection and budget can both be right and the failover still fails, because the standby region has a missing secret, an unreplicated role, or a model version behind primary. This is Module 11’s “untested failover” in its AI-serving form, and it is only ever found by failing over.
Scaling
Section titled “Scaling”The failover decision itself is cheap and does not need to scale; what scales is the cost of the stages after it. Warm-up grows with model size and with the number of replicas that need weights at once — a standby loading many replicas in parallel is contending for the same storage throughput, so the per-replica figure in a budget is optimistic unless it was measured under that concurrency.
More regions do not shorten detection. They add choices about where to fail over to, and each standby has its own tier and its own warm-up, so a multi-region plan needs a budget per target, not one for the platform.
Security
Section titled “Security”A failover moves the security boundary along with the traffic. The standby has to hold the same credentials, network policy, and model-artifact access as primary, provisioned through the same path — a role or a secret added to primary by hand is exactly the drift that surfaces mid-incident. The controller is also a high-value target: anything that can feed it unhealthy probes can trigger a failover, so the health signal needs the same integrity as any other control-plane input.
Trade-offs
Section titled “Trade-offs”Patience vs. false failovers
How long a trigger waits sets both of its numbers at once. In the backing lab, 3 consecutive failed probes detected in 20 seconds and failed over falsely 25 times a day; 12 detected in 110 seconds and failed over falsely 0.2 times a day. There is no setting that is both fast and quiet, so the choice is about which cost the platform can absorb — and the magnitudes move with the blip distribution, so re-fit them from real probe history rather than adopting someone else’s.
Window rules vs. consecutive-failure rules
Less different than they look. A window of W seconds behaves like W / interval + 1 consecutive
failures, and in the lab the two families interleave in one ordering by how long they wait. Choose
the one your team will reason about correctly, and spend the design effort on how long it waits
and on what it does with silence.
Standby tier vs. cost
A hot standby removes warm-up and costs a second serving fleet. A cold one is cheap and, with a large model, may make the RTO impossible regardless of the trigger. That makes the standby tier the first RTO decision, not the last: pick the tier that leaves room for detection, then tune detection inside the room it leaves.
The standby tier is where the money goes. Hot means paying for GPU capacity that serves nothing until the day it serves everything; warm means paying for capacity without the model loaded; cold means paying for neither and accepting the longest warm-up. False failovers have a cost too, less visible on a bill: each one is engineering time, a cold cache in the region now serving, and the risk of a failover going wrong. A trigger tuned to save detection seconds by firing on blips is buying those seconds with that cost, many times a day.
Observability
Section titled “Observability”- Every stage of every failover, timed. Detection, promotion, shift, and warm-up as separate spans, so a missed RTO says which stage spent the budget.
- Probe arrivals, not just probe results. The silent-region failure is invisible to a dashboard of health-check outcomes, because a missing probe is not an outcome. Alert on probes that were due and did not arrive.
- False failovers as a rate. Counted per day, with the blips that caused them, so the trigger’s patience can be re-fit from evidence.
- Warm-up measured, not assumed. Weight-load time per replica under the concurrency of a real failover, recorded during every exercise, so the budget uses a number rather than a datasheet.
Production Deployment
Section titled “Production Deployment”Before the failover runbook is trusted
- The RTO is written out as a budget, stage by stage, with each downstream stage measured on this platform.
- The detection window is derived from what the budget leaves, not set equal to the RTO.
- The standby tier leaves positive room for detection at the current model size — and the budget is redone whenever the model grows.
- A probe that is due and does not arrive counts as a failure, and that is covered by a test.
- The controller does not fire until it has a minimum history after a restart.
- The false-failover rate is measured from real probe history and owned by someone.
- Standby configuration is provisioned through the same path as primary, and diffed.
- Failover is exercised on a schedule, end to end, with each stage timed.
What the backing lab does not measure
RPO. Replication lag at the moment of failure is Module 11’s third failure mode, and the lab deliberately leaves it out: an honest simulation needs a model of how lag behaves under the load that precedes an outage, and any such model would assume the answer. A failover that meets its RTO can still lose more data than the RPO allows, and that has to be measured separately.
Hands-on Lab
A seeded simulation of regions, probes, blips, and outages, with Module 11’s window controller and a consecutive-failure trigger, each with and without a silence rule. It measures false failovers per day against detection delay, sets detection against a recovery budget for hot, warm, and cold standbys, and pins the restarted-controller flaw in a test. Read the lab documentation →
labs/region-failover-budgetproduction-shaped
Interview Questions
Section titled “Interview Questions”Your RTO is 15 minutes. How long should the failover trigger wait?
Not 15 minutes. Write the budget out first: promotion, traffic shift, and — for model serving — loading weights onto the standby’s GPUs all happen after detection. The trigger gets what is left. If what is left is negative, the answer is a warmer standby or a smaller model, not a faster trigger.
A region goes completely dark and your controller never fails over. Why?
Because it reasons only about the checks it receives, and a dark region sends none. “No unhealthy checks” and “no checks” are different states, and a controller that collapses them reads silence as health. The fix is to count a probe that was due and did not arrive as a failure — and to alert on missing probes, because no dashboard of check results will show the absence of one.
How do you choose between detecting fast and not failing over on blips?
By measuring both from real probe history and choosing which cost to pay. Patience trades one for the other directly: fewer false failovers always means later detection. The useful framing is that a false failover is an incident too, so the comparison is incidents caused against RTO risk taken — and the trigger’s patience is a policy with an owner, re-fit as traffic changes.
Why is failing over AI serving different from failing over a web tier?
Warm-up. A stateless web tier in a standby region can take traffic as soon as it is routed there. A model server needs the model in GPU memory first, and at large model sizes that is minutes — more if capacity has to be found. It makes standby tier and model size first-order RTO decisions, and it means the budget has to be redone whenever the model grows.
A deploy restarts your failover controller and it immediately fails the region over. What happened?
A window trigger with no history treats one unhealthy check as every check in its window. Restarts happen during deploys, when probes are noisiest, so the first bad probe after the restart fires it. The fix is to arm the trigger only after it has a minimum history — the same reason the window exists in the first place.