Skip to content

SLO-Driven AI Operations

Listen to this page4:31
Read the transcript

1. Burn rate as a spendable, alertable number

Host: So let’s start with a number that most dashboards get wrong: the error budget. The instinct is to clamp it at zero once you’ve blown it, right? Zero means bad, done, move on. But this lab does something different — it lets that number go negative.

Guest: Right, and that’s the whole point. Budget remaining fraction going negative tells you how far over you are, not just that you’re over. Clamping to zero erases that distinction exactly when you need it most. It reframes reliability as something you spend deliberately, not a target you either hit or miss.

Host: Okay, so if we’re treating it as a spendable quantity, the next problem is alerting on the rate you’re spending it — and that’s where it gets tricky, because a single window either pages you on noise or catches things too late.

Guest: Exactly, that’s a real trade-off, not just a tuning knob. A short window is twitchy — one bad minute and you’re paged for nothing. A long window smooths that out but by the time it crosses threshold you’ve already burned way more budget than you should have. So the lab requires both: a long window like an hour to confirm this is a real burn, and a short window like five minutes riding alongside it so the alert actually clears within minutes once the burn stops. Neither window alone can do both jobs.

2. One severity, two consumers: the cooldown override and escalation over retry

Host: So once you’ve got that severity number, what actually consumes it? You mentioned it drives the autoscaler directly — walk me through that.

Guest: Right, so normally the autoscaler has a cooldown timer to stop it flapping on noisy utilization samples — you scale up, then you wait before scaling again. But the controller checks burn severity before it even looks at cooldown. If severity is fast-burn and you’re under max replicas, it scales up immediately, cooldown or not.

Host: That feels risky on paper — bypassing your own safety mechanism. Why is that actually the right call instead of just a bug waiting to happen?

Guest: Because a cooldown is designed to suppress noise, and a page-severity burn isn’t noise — it’s the service actively failing users right now. Waiting out a timer built for noisy metrics is the wrong response to a real outage. It’s a narrow escape hatch for exactly one case; every other decision still respects cooldown normally. And the same severity number does this on the incident side too — the runbook’s three steps, restart, scale up, page on-call, each run under their own timeout, and if a step doesn’t resolve things it escalates to a named contact instead of retrying blindly. If all three steps burn through without resolution, the incident reports EXHAUSTED with the last contact paged, rather than looping forever or just quietly giving up.

3. Trying it yourself and what’s still a stand-in

Host: So if someone wants to see all this actually run, what’s the fastest path? I imagine there’s a way to spin it up locally and just watch the severity number change behavior in real time.

Guest: Yeah, it’s a small FastAPI app — go into the labs slo-driven-ai-operations directory, pip install with the dev extras, and run uvicorn. It registers one demo service, checkout-api, with a real runbook of restart, scale up, page on-call. You post a healthy autoscale request first and see ordinary proportional scaling, then you hammer it with fifty failed requests in a loop, and now the SLO endpoint reports page severity and the incident you open escalates the same way we described. Run pytest after — 32 tests, plus ruff and mypy — so you’re not just trusting the demo, you’re watching the same invariants get checked in code.

Host: Before we wrap — is there anything here that’s a stand-in, something you wouldn’t ship as-is?

Guest: One honest gap: the Clock is injected and deterministic, which is great for tests but obviously not talking to a real metrics backend yet. The important thing is that it’s the exact seam where one would plug in a real metrics backend, so wiring in real burn-rate data is an integration, not a rewrite of the severity or runbook logic. That’s really the whole point of this lab: get the one severity computation right and boring, and everything downstream, alerting, scaling, incidents, just inherits that correctness.

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

Three mechanisms that turn “is the platform healthy” from a judgement call into a control loop: error-budget tracking with multiwindow burn-rate alerting, an autoscaling controller that reacts differently to a fast burn than to ordinary load, and runbook execution that escalates rather than retrying. All three read severity from the same SLOTracker.report() call, so they cannot disagree about how bad things currently are.

Source: labs/slo-driven-ai-operations

  • Multiwindow, multi-burn-rate alerting. A rule fires only when a long window (e.g. 1h) and a short window (e.g. 5m) both exceed its threshold. The long window filters out blips; the short window lets the alert clear within minutes of the burn stopping.
  • Error budget as a spendable quantity. budget_remaining_fraction goes negative once a service has overspent, which is more useful than clamping at zero — it says how far over, not merely that it is over.
  • An SLO-aware autoscaler. The proportional (HPA-style) target and cooldown behave like an ordinary autoscaler, but a page-severity burn bypasses the cooldown entirely.
  • Escalation instead of retry. Each runbook step is bounded by its own asyncio.timeout, and an unresolved step escalates to a named contact before the next step runs.
  • An honest terminal state. When every step is exhausted without resolution, the incident reports EXHAUSTED with the last contact paged, rather than looping or silently stopping.
  • Deterministic time. An injectable Clock means burn-rate windows and cooldowns are tested without a single real sleep — the whole suite runs in under a second.

A cooldown exists to stop an autoscaler flapping on noisy utilization samples. But a burn rate over the page threshold is not noise — it means the service is failing users right now, and waiting out a timer designed to suppress noise is the wrong response. The controller therefore checks burn severity before it checks cooldown:

def decide(self, *, current_replicas, utilization, alert_severity):
if (alert_severity == self._policy.fast_burn_severity
and current_replicas < self._policy.max_replicas):
# an active outage is not something to wait out
return self._clamped_decision(current_replicas + 1, current_replicas, ...)
if self._in_cooldown():
return ScalingDecision(HOLD, current_replicas, "within cooldown window")
...

Every other scaling decision still respects the cooldown. The override is a deliberate, narrow escape hatch for the one case the cooldown was never meant to slow down.

Terminal window
cd labs/slo-driven-ai-operations
python3.12 -m venv .venv
source .venv/bin/activate
pip install -e '.[dev]'
uvicorn platform_ops.app:app --reload

The app registers one demo service, checkout-api, targeting 99.9% success over a 30-day budget window, with a three-step runbook: restart_service → scale_up_capacity → page_on_call.

Record outcomes, then watch severity change what the autoscaler does:

Terminal window
# a healthy service: ordinary proportional scaling
curl -s -X POST localhost:8000/v1/services/checkout-api/autoscale \
-H 'content-type: application/json' -d '{"current_replicas": 4, "utilization": 0.9}'
# now burn the budget hard
for _ in $(seq 50); do
curl -s -X POST localhost:8000/v1/services/checkout-api/requests \
-H 'content-type: application/json' -d '{"success": false}' > /dev/null
done
curl -s localhost:8000/v1/services/checkout-api/slo # severity: page
curl -s -X POST localhost:8000/v1/incidents \
-H 'content-type: application/json' -d '{"service": "checkout-api"}'

The incident’s severity is derived from the SLO report, so the same failures that changed the scaling decision also change how far the runbook escalates.

Terminal window
pytest # 32 tests
ruff check .
mypy src
  1. An error budget reframes reliability as something you spend rather than a target you hit or miss — that framing is what lets a team trade reliability for velocity deliberately instead of accidentally.
  2. Multiwindow, multi-burn-rate alerting exists to escape a trade-off a single-window threshold cannot: short windows page on noise, long windows react too late. Being able to state that precisely matters more in an interview than naming the technique.
  3. Deriving the scaling decision and the incident’s severity from one report() call means the two subsystems cannot disagree — a common source of contradictory automation during real incidents.
  4. EXHAUSTED is a feature. Incident automation that hides “I am out of steps and a human owns this now” withholds the first thing an on-call engineer needs to know.
  5. The injectable Clock is not only a testing convenience — it is the seam where a real metrics backend substitutes in, which is why the production gap above is an integration rather than a rewrite.