SLO-Driven AI Operations
Read the transcript
1. Burn rate as a spendable, alertable number
Host: So let’s start with a number that most dashboards get wrong: the error budget. The instinct is to clamp it at zero once you’ve blown it, right? Zero means bad, done, move on. But this lab does something different — it lets that number go negative.
Guest: Right, and that’s the whole point. Budget remaining fraction going negative tells you how far over you are, not just that you’re over. Clamping to zero erases that distinction exactly when you need it most. It reframes reliability as something you spend deliberately, not a target you either hit or miss.
Host: Okay, so if we’re treating it as a spendable quantity, the next problem is alerting on the rate you’re spending it — and that’s where it gets tricky, because a single window either pages you on noise or catches things too late.
Guest: Exactly, that’s a real trade-off, not just a tuning knob. A short window is twitchy — one bad minute and you’re paged for nothing. A long window smooths that out but by the time it crosses threshold you’ve already burned way more budget than you should have. So the lab requires both: a long window like an hour to confirm this is a real burn, and a short window like five minutes riding alongside it so the alert actually clears within minutes once the burn stops. Neither window alone can do both jobs.
2. One severity, two consumers: the cooldown override and escalation over retry
Host: So once you’ve got that severity number, what actually consumes it? You mentioned it drives the autoscaler directly — walk me through that.
Guest: Right, so normally the autoscaler has a cooldown timer to stop it flapping on noisy utilization samples — you scale up, then you wait before scaling again. But the controller checks burn severity before it even looks at cooldown. If severity is fast-burn and you’re under max replicas, it scales up immediately, cooldown or not.
Host: That feels risky on paper — bypassing your own safety mechanism. Why is that actually the right call instead of just a bug waiting to happen?
Guest: Because a cooldown is designed to suppress noise, and a page-severity burn isn’t noise — it’s the service actively failing users right now. Waiting out a timer built for noisy metrics is the wrong response to a real outage. It’s a narrow escape hatch for exactly one case; every other decision still respects cooldown normally. And the same severity number does this on the incident side too — the runbook’s three steps, restart, scale up, page on-call, each run under their own timeout, and if a step doesn’t resolve things it escalates to a named contact instead of retrying blindly. If all three steps burn through without resolution, the incident reports EXHAUSTED with the last contact paged, rather than looping forever or just quietly giving up.
3. Trying it yourself and what’s still a stand-in
Host: So if someone wants to see all this actually run, what’s the fastest path? I imagine there’s a way to spin it up locally and just watch the severity number change behavior in real time.
Guest: Yeah, it’s a small FastAPI app — go into the labs slo-driven-ai-operations directory, pip install with the dev extras, and run uvicorn. It registers one demo service, checkout-api, with a real runbook of restart, scale up, page on-call. You post a healthy autoscale request first and see ordinary proportional scaling, then you hammer it with fifty failed requests in a loop, and now the SLO endpoint reports page severity and the incident you open escalates the same way we described. Run pytest after — 32 tests, plus ruff and mypy — so you’re not just trusting the demo, you’re watching the same invariants get checked in code.
Host: Before we wrap — is there anything here that’s a stand-in, something you wouldn’t ship as-is?
Guest: One honest gap: the Clock is injected and deterministic, which is great for tests but obviously not talking to a real metrics backend yet. The important thing is that it’s the exact seam where one would plug in a real metrics backend, so wiring in real burn-rate data is an integration, not a rewrite of the severity or runbook logic. That’s really the whole point of this lab: get the one severity computation right and boring, and everything downstream, alerting, scaling, incidents, just inherits that correctness.
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Three mechanisms that turn “is the platform healthy” from a judgement call into a control loop:
error-budget tracking with multiwindow burn-rate alerting, an autoscaling controller that reacts
differently to a fast burn than to ordinary load, and runbook execution that escalates rather than
retrying. All three read severity from the same SLOTracker.report() call, so they cannot disagree
about how bad things currently are.
Source: labs/slo-driven-ai-operations
Control loop
Section titled “Control loop”flowchart TB
R["Request outcomes
(success / failure)"] --> T
subgraph T["SLOTracker"]
W["Rolling window,
pruned to the SLO's
budget window"]
B["Burn rate =
observed error rate
÷ error budget"]
W --> B
end
B --> LW["Long window
(e.g. 1h)"]
B --> SW["Short window
(e.g. 5m)"]
LW --> AND{"both over
threshold?"}
SW --> AND
AND -->|no| OK["severity: ok"]
AND -->|yes| SEV["severity:
ticket / page"]
OK --> AC
SEV --> AC
SEV --> IR
subgraph AC["AutoscalingController"]
FB{"severity
== page?"}
FB -->|"yes — skip cooldown,
an outage is not
something to wait out"| UP["scale up"]
FB -->|no| CD{"in cooldown?"}
CD -->|yes| HOLD["hold"]
CD -->|no| PROP["HPA-style
proportional target,
clamped to bounds"]
end
subgraph IR["IncidentRunner"]
S1["step 1
bounded by timeout"] -->|unresolved| E1["escalate to
named contact"]
E1 --> S2["step 2"]
S2 -->|unresolved| E2["escalate"]
E2 --> S3["... until resolved
or EXHAUSTED"]
endWhat it demonstrates
Section titled “What it demonstrates”- Multiwindow, multi-burn-rate alerting. A rule fires only when a long window (e.g. 1h) and a short window (e.g. 5m) both exceed its threshold. The long window filters out blips; the short window lets the alert clear within minutes of the burn stopping.
- Error budget as a spendable quantity.
budget_remaining_fractiongoes negative once a service has overspent, which is more useful than clamping at zero — it says how far over, not merely that it is over. - An SLO-aware autoscaler. The proportional (HPA-style) target and cooldown behave like an ordinary autoscaler, but a page-severity burn bypasses the cooldown entirely.
- Escalation instead of retry. Each runbook step is bounded by its own
asyncio.timeout, and an unresolved step escalates to a named contact before the next step runs. - An honest terminal state. When every step is exhausted without resolution, the incident
reports
EXHAUSTEDwith the last contact paged, rather than looping or silently stopping. - Deterministic time. An injectable
Clockmeans burn-rate windows and cooldowns are tested without a single real sleep — the whole suite runs in under a second.
Why a fast burn overrides the cooldown
Section titled “Why a fast burn overrides the cooldown”A cooldown exists to stop an autoscaler flapping on noisy utilization samples. But a burn rate over the page threshold is not noise — it means the service is failing users right now, and waiting out a timer designed to suppress noise is the wrong response. The controller therefore checks burn severity before it checks cooldown:
def decide(self, *, current_replicas, utilization, alert_severity): if (alert_severity == self._policy.fast_burn_severity and current_replicas < self._policy.max_replicas): # an active outage is not something to wait out return self._clamped_decision(current_replicas + 1, current_replicas, ...)
if self._in_cooldown(): return ScalingDecision(HOLD, current_replicas, "within cooldown window") ...Every other scaling decision still respects the cooldown. The override is a deliberate, narrow escape hatch for the one case the cooldown was never meant to slow down.
Run it
Section titled “Run it”cd labs/slo-driven-ai-operationspython3.12 -m venv .venvsource .venv/bin/activatepip install -e '.[dev]'uvicorn platform_ops.app:app --reloadThe app registers one demo service, checkout-api, targeting 99.9% success over a 30-day budget
window, with a three-step runbook: restart_service → scale_up_capacity → page_on_call.
Record outcomes, then watch severity change what the autoscaler does:
# a healthy service: ordinary proportional scalingcurl -s -X POST localhost:8000/v1/services/checkout-api/autoscale \ -H 'content-type: application/json' -d '{"current_replicas": 4, "utilization": 0.9}'
# now burn the budget hardfor _ in $(seq 50); do curl -s -X POST localhost:8000/v1/services/checkout-api/requests \ -H 'content-type: application/json' -d '{"success": false}' > /dev/nulldone
curl -s localhost:8000/v1/services/checkout-api/slo # severity: pagecurl -s -X POST localhost:8000/v1/incidents \ -H 'content-type: application/json' -d '{"service": "checkout-api"}'The incident’s severity is derived from the SLO report, so the same failures that changed the scaling decision also change how far the runbook escalates.
Verify it
Section titled “Verify it”pytest # 32 testsruff check .mypy srcPrincipal-level discussion points
Section titled “Principal-level discussion points”- An error budget reframes reliability as something you spend rather than a target you hit or miss — that framing is what lets a team trade reliability for velocity deliberately instead of accidentally.
- Multiwindow, multi-burn-rate alerting exists to escape a trade-off a single-window threshold cannot: short windows page on noise, long windows react too late. Being able to state that precisely matters more in an interview than naming the technique.
- Deriving the scaling decision and the incident’s severity from one
report()call means the two subsystems cannot disagree — a common source of contradictory automation during real incidents. EXHAUSTEDis a feature. Incident automation that hides “I am out of steps and a human owns this now” withholds the first thing an on-call engineer needs to know.- The injectable
Clockis not only a testing convenience — it is the seam where a real metrics backend substitutes in, which is why the production gap above is an integration rather than a rewrite.
Related
Section titled “Related”- Architecture: AI Reliability Platform — the design-review companion: problem framing, constraints, cost, observability, and a pre-traffic checklist.
- Module 12: Observability — SLOs, error budgets, and the signal pipeline this lab consumes.
- Module 10: Kubernetes — the scheduling and scaling primitives a real autoscaling target would drive.
labs/async-ai-gateway— theproduction-readyreference this lab is measured against.