Skip to content

Module 12: Observability

Listen to this page14:46
Read the transcript

1. The one question observability has to answer fast

Host: Welcome back. This module is about observability, and I want to start with the question that actually matters when something’s on fire: what happened to this one request, and why did it happen? Not the aggregate, not the dashboard trend — this specific request, right now.

Guest: Right, and that’s the framing we should hold onto for the whole module. Metrics tell you something’s wrong across the system — error rate ticked up, latency’s spiking. Logs and traces are what let you zoom into one request and reconstruct exactly what it did and why it failed.

Host: So those three feel like the whole observability toolkit — until you’re running an AI system, apparently. You mentioned there’s a fourth signal that has to sit alongside them. Why can’t the classic three just cover it?

Guest: Because a slow, cheap, confidently wrong answer looks completely normal on every one of those dashboards — fast response, low latency, no error thrown. Metrics, logs, and traces are structurally blind to whether the answer was actually correct, which is exactly why evaluation-based quality monitoring has to be its own independently watched signal.

2. Three signals, one shared identifier — and their common blind spot

Host: So walk me through the actual mental model here, because ‘observability’ gets thrown around like it’s one thing. It’s really three separate signals, right?

Guest: Right — metrics, logs, and traces, and they’re stitched together by a shared request_id. Metrics are the aggregated numbers, request rate, error rate, latency percentiles, cheap to store and great for spotting that something’s wrong right now. Logs are the discrete timestamped events, structured key-value ideally, and that request_id is what lets you pull every log line across every service that touched one specific request. Traces are the tree of timed spans showing the actual path that request took, every downstream call, in sequence — it’s the only one of the three that shows you causality directly.

Host: Okay, so between the three of them you can reconstruct exactly what a request did, end to end. But you just said none of that tells you if the answer was right.

Guest: Exactly — all three describe how a request executed, not whether the output was correct. A hallucinating model produces a 200, normal latency, a clean trace. There’s no field in any of those three signals for ‘this answer was confidently wrong,’ which is why that evaluation harness has to run as its own independently watched layer wrapping the whole system, not as a feature bolted onto metrics or logs.

3. Why agent fan-out demands span-level tracing

Host: So let’s zoom into that trace itself, because a single user request to an agent isn’t one call anymore — it’s fanning out into multiple model calls and tool invocations, right, the stuff we covered back in the agent engineering and MCP modules.

Guest: Right, and that’s exactly why span-level tracing matters specifically for agentic systems. You’ve got a top-level request span, and nested underneath it you get llm.call spans and tool.call spans. Each of those is its own span, not just one blob of latency for the whole request.

Host: And without that nesting, what do you actually see when something inside goes slow?

Guest: Just elevated latency on the outermost span, full stop. You have no idea which inner step caused it — the outer span can’t tell you which step is responsible, only that something did.

4. RED for services, USE for the resources underneath them

Host: So if the trace tells you something’s wrong but not which inner step, what do engineers actually pull up first when they get paged? There’s got to be a standard starting point before you even go span-diving.

Guest: Yeah, that’s RED — Rate, Errors, Duration, usually a latency histogram. It’s the metric set for the service itself, and it maps onto the exact questions an on-call engineer asks first: are we getting hit harder than usual, are requests failing, are they slow. But RED only tells you the symptom is happening, not why.

Guest: That’s where USE comes in — Utilization, Saturation, Errors — but for the resources underneath the service: CPU, memory, GPU, the KV cache budget we covered back in Module 9. RED shows you the symptom, and USE is usually what explains why it’s happening.

5. SLIs, SLOs, and spending an error budget on purpose

Host: So once you know CPU or KV cache saturation is the culprit, how do you decide when that actually deserves a page versus just being noted for later? That’s presumably where SLOs come in.

Guest: Right, and the vocabulary matters here. An SLI is the specific measured thing, like p99 latency. The SLO is your target for it, say p99 under 800 milliseconds over a rolling 28 days. And the error budget is how much violation of that target you can afford before it’s genuinely unacceptable. The trick is alerting not on every threshold crossing but on burn rate — how fast you’re consuming that budget — because a brief latency blip and a trend that will exhaust your whole month’s budget by Tuesday look identical on a threshold alert, but they demand completely different responses.

Host: And that’s not just theory — the SLO-driven-ops lab actually builds a controller around that distinction, doesn’t it?

Guest: Exactly, it implements multiwindow burn-rate alerting with an autoscaler that treats a fast burn as categorically different from ordinary load. Normally the controller respects a cooldown so it doesn’t flap on noisy utilization samples, but the code explicitly checks burn severity before it checks cooldown at all — if it’s a fast burn, it scales up immediately, because an active outage isn’t something you wait out a noise-suppression timer for. Every other scaling decision still obeys the cooldown; this is a narrow, deliberate escape hatch for the one case that timer was never meant to slow down.

6. Where async code silently breaks a trace

Host: Let’s shift to something that trips up even teams who’ve done tracing right everywhere else: async code. We flagged this back in Module 1 as a landmine for implicit state — does it come back to bite tracing specifically?

Guest: Constantly, and it’s because a trace only works if the trace ID and parent span ID physically travel with the request across every boundary — an HTTP hop, a queue message, a spawned task. The moment you fire off a background task without explicitly carrying that context forward, the new span has no parent to attach to.

Host: And what does that look like to the person debugging it — do they at least get an error pointing at the gap?

Guest: That’s the cruel part, no. You just get an orphan span, or a trace that looks incomplete with a step missing, and nothing throws — the background task ran fine, it just isn’t linked. It’s a silent structural gap, not a failure, so it’s exactly the kind of thing that hides until someone’s staring at a trace during an incident wondering where a step went.

7. Structured logs, AI-specific fields, and redact-before-storage

Host: Okay, so let’s go from the tracing gap to the logs themselves. Why is a structured log line actually different from just writing a good sentence to a file?

Guest: Because during an incident nobody has time for regex archaeology. If every log is a JSON object with consistent fields — request_id, tenant_id, latency_ms — you can query it instantly, filter it, join it with the trace. Free text means someone’s grepping and guessing at 2am, and for AI systems you want token counts, model identifiers, and cost sitting as structured fields on that same line, not stashed in a separate billing system you have to manually correlate after the fact.

Host: And prompts and responses themselves — you’d think logging the full content is exactly what you want for debugging, so where’s the catch?

Guest: The catch is those prompts and responses carry the same sensitive data everything else is built to protect — pasted credentials, customer PII, whatever a user typed in a support query. So redaction has to happen before it ever reaches the logging pipeline, not as an access-control rule on who’s allowed to query the logs afterward. Once it’s stored unredacted, restricting query access doesn’t undo the fact that you already have it sitting there — and the telemetry store holding all this is now its own attack surface, not just a debugging convenience.

8. A minimal tracer: cost, tokens, and redaction as span attributes

Host: So let’s make that redact-before-storage promise concrete. Walk me through what this tracer actually looks like in code.

Guest: It’s small on purpose. You’ve got a Span that’s just a name plus a dictionary of attributes, and a Tracer whose start_span context manager opens a span, times it, and appends it to a list of finished spans when it closes. The interesting part is inside call_model — before anything gets set on the span, the prompt goes through redact, which does a straight substring replace of anything sensitive with a literal ‘[redacted]’ marker. Only the redacted version ever gets assigned to llm.prompt_redacted, alongside llm.model, token counts, and an estimated cost computed straight from the token count. So one span attribute set gives you what happened, how much it cost, and what model did it — no separate cost system to join against.

Host: And the lab makes you prove that redaction actually held, not just that the function exists somewhere.

Guest: Right — the test calls call_model with a prompt containing a marker like ‘api_key: sk-test-123’ in the sensitive_substrings list, then it doesn’t check the return value, it inspects tracer.finished_spans and asserts that marker string doesn’t appear in any attribute on any span. That’s a materially different claim than ‘we have a redact function’ — it’s checking the artifact that actually gets persisted and queried later.

9. Case study: chasing a latency regression into one tool call

Host: Okay, let’s put all of this together with an actual incident. The p99 latency SLO starts burning fast, error budget’s draining, on-call gets paged. What do the RED metrics actually show you at that point?

Guest: Just elevated duration on the top-level request span — that’s it. RED is service-level by design, so all you know is ‘requests are slow,’ which in an agentic system could mean the model’s slow, a tool’s slow, or the orchestration logic itself is stalling. That ambiguity alone isn’t enough to diagnose anything.

Host: Which is exactly the wrong move here, based on what you’re setting up.

Guest: Right — that’s the trap. Once you drop into the nested spans, llm.call and tool.call under that same trace, the picture changes completely: the model calls are all executing at normal latency, and the slowdown is isolated to one specific tool call wrapping a downstream API that’s degraded. Without span-level tracing through that fan-out, this reads as a generic model latency regression and sends the on-call engineer to debug the wrong system entirely — the SLO told them something was wrong, but only the trace tree told them where.

10. What goes wrong: cardinality explosions, alert fatigue, and unbounded trace cost

Host: That case study worked out because someone could actually query the trace tree fast and find the one bad tool call. But all this instrumentation isn’t free — where does it actually bite you operationally?

Guest: The classic one is cardinality. Someone adds a raw user ID or request ID as a metric label instead of a bounded dimension like tenant tier, and now the backend is storing a distinct time series per user. Nobody notices at first — it’s silent until query performance or storage cost degrades enough that someone goes looking, and by then it’s hard to trace back to the label that caused it.

Host: And that’s not just a cost problem for the team that added it, right? It slows the whole backend down for everyone.

Guest: Exactly, it degrades queries for every team hitting that store. And it compounds with two other cost dynamics: alerting on every threshold crossing instead of burn rate trains on-call to ignore pages, and trace and log volume scale with request volume, not with usefulness — the millionth identical successful trace tells you nothing, which is why sampling and tiered logging exist. Multi-region setups add a fourth wrinkle too — a single global backend every region writes to synchronously just reimports the cross-region coupling Module 11 argues against everywhere else in the architecture.

11. The fourth signal: why evaluation has to run on its own

Host: So let’s land on the thing this whole module has been circling. A model provider quietly swaps the version behind a stable endpoint, or someone edits a production prompt, and every dashboard we’ve built stays green — the trace looks clean, latency is fine, the RED metrics don’t move. Nothing here catches that.

Guest: Right, because all three signals describe execution, not correctness. Status code, duration, path through the system — none of that encodes whether the answer was actually good. That’s why Module 4’s evaluation harness has to run continuously and independently, scoring live or sampled output against a rubric on its own schedule, not triggered by anything observability noticed, because observability structurally can’t notice. Treat a model or prompt change exactly like a deploy: canary it, compare eval scores against baseline, keep a rollback path — the same rigor as code, because it changes behavior the same way code does.

Host: And one last thing before we close — all this telemetry we’ve spent the module building is itself sensitive. As we covered earlier, redact before storage, not after — and beyond that, the traces and logs need real access control, because a database full of request payloads and internal topology is a target, not just a debugging convenience. That’s the module: three fast signals for what happened, one independent signal for whether it should have happened at all, and don’t forget to lock the door on all four.

Not covered

The planner wanted these and found nothing in the source to support them:

  • A step-by-step OpenTelemetry SDK instrumentation walkthrough
  • A head-to-head comparison of specific commercial tracing/logging backend vendors
  • Detailed cost modeling of running a self-hosted metrics/trace backend at scale

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

Observability exists to answer one question fast during an incident: what actually happened to this request, and why. Metrics tell you something is wrong in aggregate; logs and traces tell you what happened to a specific request; and for AI systems specifically, a fourth signal — evaluation- based quality monitoring — tells you the system is answering correctly at all, a failure mode the other three are structurally blind to, since a slow, cheap, confidently wrong answer looks identical to a correct one on every latency or error-rate dashboard.

Three signals, correlated by a shared identifier, cover most production debugging:

  • Metrics — aggregated numbers over time (request rate, error rate, latency percentiles). Cheap to store at high cardinality-collapsed granularity; good for “is something wrong right now,” bad for “what happened to this one request.”
  • Logs — discrete, timestamped events, ideally structured (key-value, not free text) and carrying a request_id that ties them to the same request across every service it touched.
  • Traces — a tree of timed spans representing one request’s actual path through the system, including every downstream call it made — the only signal that shows causality and sequencing directly.

A model can be fast, cheap, and wrong, and none of your dashboards will show it

Metrics, logs, and traces all describe how a request executed — latency, errors, the call graph — none of them describe whether the answer was correct. An LLM that hallucinates confidently produces a 200 response, a normal latency, and a clean trace. This is exactly why Module 4’s evaluation harness has to run as its own, separately monitored signal — observability in the traditional sense cannot substitute for it, no matter how much telemetry you add.

The nested spans under llm.call and tool.call are what make a trace useful for an agentic system specifically — a single top-level request can fan out into several model calls and tool invocations (per Module 5 and Module 6), and without span-level tracing through that fan-out, a slow or failing step is invisible except as elevated latency on the outermost span, with no indication of which inner step actually caused it.

RED and USE. For request-serving services, the RED method (Rate, Errors, Duration, usually as a latency histogram) is the standard starting metric set — it maps directly onto the questions an on-call engineer actually asks first. The USE method (Utilization, Saturation, Errors) is the complementary view for the resources underneath a service — CPU, memory, GPU, the KV cache budget Module 9 covers — and it’s what usually explains why a RED-method symptom is happening.

SLIs, SLOs, and error budgets. A Service Level Indicator is a specific measured metric (e.g., p99 latency); a Service Level Objective is the target for it (e.g., p99 < 800ms over a rolling 28 days); the error budget is the amount of SLO violation you can afford before it becomes unacceptable. Alerting on SLO burn rate — how fast the error budget is being consumed — produces fewer, more actionable pages than alerting on every individual metric threshold crossing, because it distinguishes a brief blip from a trend that will actually exhaust the budget.

Distributed tracing and context propagation. A trace is only as useful as its context propagation — the trace ID and parent span ID have to travel with the request across every process boundary (an HTTP call, a queue message, an async task) for spans from different services to link into one coherent trace. This is exactly where async code (per Module 1) commonly breaks tracing silently: a fire-and-forget background task started without explicitly propagating the current trace context produces an orphan span with no parent, and the trace looks incomplete with no error to indicate why.

Structured logging. A log line as a JSON object with consistent field names (request_id, tenant_id, latency_ms) is queryable; a log line as free-form text requires regex archaeology during an incident, exactly when time matters most. The specific discipline for AI systems: log token counts, model identifiers, and cost as structured fields on the same log line as the request outcome, not in a separate system that has to be manually correlated afterward.

Prompt and response logging vs. redaction. Logging full prompts and responses is often essential for debugging quality issues, but prompts and responses can contain the same sensitive data the rest of the system is built to protect — customer PII, credentials pasted into a support query. Redaction has to happen before data reaches the logging pipeline, not as a policy about who’s allowed to query the logs afterward, which doesn’t undo the exposure of having stored it unredacted in the first place.

Research Note

The chapters that formalized RED-style metrics, SLIs/SLOs, and error budgets as the standard vocabulary for production monitoring — worth reading directly since the SLO/error-budget framing is precise in a way that’s easy to blur in casual use.

Source: Google SRE Book, "Monitoring Distributed Systems" and "Service Level Objectives"

The pipeline that guards your system is a system too, and green is its dashboard. Everything above treats production telemetry as the thing that can mislead you. The delivery pipeline earns the same suspicion for the same structural reason: a pipeline that ran its tests and a pipeline that skipped them emit an identical signal, and that signal is a green check.

A red build is looked at within minutes, because someone is blocked. A false green is looked at by nobody, because nothing is blocked — it produces confidence that was never earned, which is worse than no confidence at all. That asymmetry is the whole problem: the outcome most in need of scrutiny is the one wearing the colour of success.

So CI correctness is not “the command exited 0”. It is four separate claims, and the exit code establishes only the first:

CI correctness = the check succeeded
+ the check actually executed
+ it covered the scope you believed it covered
+ unexpected skips were surfaced rather than swallowed

Exit code 0 proves that the process it invoked considered itself successful. It does not prove that any tests were discovered, that the typechecker was aimed at the package you changed, that the linter saw the directory you added, or that the integration suite did anything beyond skipping every case because a dependency was unreachable.

Why this sharpens under AI-assisted development. Module 14 argues that a check which has never failed is indistinguishable from a check that cannot fail. An agent working in a repository can change test discovery, a path filter, a workflow condition, or a config default as a side effect of the change it was actually asked to make — and it will then report success on the evidence CI handed it. The contract worth holding an agent to is not implementation changed → CI green:

implementation changed
→ the expected validation suites were discovered
→ the expected scope was executed
→ the required gates passed
→ CI green

The generalisation is validate the validator. When a change touches both an implementation and the mechanism that checks it, those are not one change, and reviewing them as one is how a weakened gate ships in the same commit as the code it existed to catch. For agentic workflows the evaluation has two questions, not one: did the proposed change pass validation, and does the validation still cover the system it claims to.

Don't assert that CI passed — assert that CI tested what you think it tested

A green build is evidence only when you can show which checks ran and over what. That is the same discipline this module applies to a latency dashboard: the number is not the property, it is a claim about the property, and the gap between them is where the incident lives. Ask any pipeline you inherit to tell you how many tests it discovered. A pipeline that cannot answer is not reporting on your system; it is reporting on itself.

Engineering Note

This handbook has shipped exactly this failure. The Async AI Gateway lab was published labelled production-ready while its CI omitted a dependency extra, so pytest’s collection aborted before a single test ran — masking thirteen type errors and a failing test behind a green check. The label was a claim, and nothing was checking the claim.

Insight sharpened by feedback from Dushyant K., who described a real false-green pipeline whose lint step exited 0 because of a configuration error, and whose eventual safeguard was to fail the build when the discovered test count fell below a known floor rather than trusting the exit code alone.

Trust has an expiry, and the same failure sits one layer up. A green build can be wrong because the expected checks never ran. A system can be just as wrong about approved because the evidence behind that approval has gone stale — the review happened, it was valid, the window lapsed, nobody owned revalidating it, and the label stayed. Both are the same defect wearing different clothes: the label no longer describes reality, and nothing in the system is required to notice. So the question a status has to answer is not only whether validation passed, but whether its evidence is still current and who is on the hook for keeping it so:

trustworthy status = validation passed
+ the expected checks executed
+ the evidence is still current
+ the control has an owner that answers

Treat production-ready, security-reviewed, evaluation-passed, human-approved, and risk-accepted as derived states backed by dated evidence, not as permanent labels a system earns once. A control worth having answers five questions: who owns it, how long its evidence is trusted, what makes it stale, what restores it, and what happens when it lapses. The last is the one usually left undefined, which is how expiry silently resolves to “carry on”.

“Who owns it” is a field that rots too. An owner recorded as a person’s name decays exactly the way the label does — someone changes team, nobody transfers the control, and the field goes on reading like an answer. Point ownership at something with defined membership instead: a role, a team, an on-call rotation. Then a departure changes who answers rather than whether anyone does, and the control does not quietly acquire a second thing that needs revalidating. Treat an owner that no longer resolves — vacant role, empty group, disbanded team — as its own staleness condition, and send it to the same expired state as lapsed evidence, because both are the case where nothing is left that is required to notice.

A resolved owner is still not a responsive one. An empty rotation resolves. So does one whose members have all moved on but whose accounts are still listed, and one whose schedule has a gap at 03:00 on a Sunday. Membership is a structural property; responsiveness is a behavioural one, and substituting the first for the second is the failure this module opened with — the number is not the property, it is a claim about the property, and the gap between them is where the incident lives. owner_resolves == true is a claim that somebody will act. It is not the acting.

Notice the shape of the last three paragraphs. The label needed a validator, the validator needed an expiry, the expiry needed an owner, and the owner now needs a liveness signal — which is itself dated evidence with an owner, so it needs one too. The tower does not terminate by adding another rung. The signals are still worth ranking, because they are not equally weak:

the group exists structural — catches nothing
members > 0 catches the emptied group
members are active accounts catches the departed-but-still-listed
someone is on shift right now catches the gapped rotation
the group acted within the window behavioural, but passive and easy to fake
an unannounced page was acknowledged the only one that is evidence

What terminates the tower is coupling, and where you cannot couple, exercise. A control that sits on a path somebody has to traverse — a deploy that will not proceed, a build that goes red — gets liveness for free, because non-response is self-announcing: the thing people wanted to do does not happen. A control whose lapse writes a row in a compliance table has no liveness at all, and needs an infinite tower of watchers to simulate it. Where coupling genuinely is not available, the only evidence of responsiveness is having exercised it — an unannounced page, a deliberately expired control, a revalidation request with a deadline and a visible consequence when it passes. This is the checklist’s “break the pipeline on purpose and confirm it goes red” applied to a human control, and Module 14’s line survives the translation intact: an owner who has never been paged is indistinguishable from an owner who cannot be reached.

Two things fall out of that. Liveness is a property of the channel, not the team — a group with an excellent page-acknowledgement record can ignore “your review expired” mail indefinitely, so test the channel the control actually uses rather than the team’s general reputation. And where you can neither couple the control nor bring yourself to exercise it, the honest move is to stop calling it a control and downgrade whatever claim rests on it, for the same reason a permanent waiver is a control that was deleted without saying so.

The cheap version needs no workflow engine — only that expiry produces a distinct state instead of leaving the old one standing:

review required every 30 days → review completed → approved
→ 30 days pass →
wrong: status stays approved
right: status becomes review-expired, and the next deploy asks for revalidation

In AI systems most controls are time-sensitive by nature. A model evaluation, a red-team pass, a prompt review, or a dataset check is evidence about a specific artifact, and the artifact moves. An approval is therefore only meaningful when bound to what was actually reviewed — the model version, the prompt version, the toolset, the policy, the eval it passed — so that changing any of them invalidates it rather than inheriting it. An agent that swaps a tool or a prompt version under an existing approval has not kept the approval; it has voided it, and the system should be able to say so without a human noticing first.

A control with no owner and no expiry eventually becomes another stale label

Don’t only validate the validator — validate that the validator is still current, and that something still answers when it fires. Fail closed where an expired control guards something irreversible, and warn where it does not; a gate that blocks everything on lapse gets routed around, and one that blocks nothing was never a gate.

Engineering Note

This handbook already runs on this principle without having stated it. Every fast-moving page declares what it was verifiedAgainst and verifiedOn, and scripts/lint-content-structure.ts fails the build once a page passes its review window — 90 days for fast-moving topics, 270 and 540 for slower ones. That is the whole pattern in one mechanism: dated evidence, a defined validity period, an owner (whoever the failing build stops), and a failure behaviour that is a red build rather than a quiet assumption that August’s verification still holds.

It was also, until this paragraph was written, missing the coupling the section above recommends. site-ci.yml has always run that linter on every pull request, but main carried no branch protection and no rulesets, so not one of those checks was actually required. The three previous versions of this page merged green because somebody chose to wait for the checks — nothing would have stopped the merge otherwise. By this module’s own checklist, a gate nothing enforces is decoration: a control that resolved, ran, and went red, whose enforcement rested entirely on whether a human still cared that week.

Checking that while writing about it is what turned the checks into required ones, admins included. The point survives the fix rather than being cancelled by it — the linter’s liveness never came from the linter, and for three versions it came from nowhere at all. It comes now from being on the path to main, which is the only place this section claims to find it.

Extended from a LinkedIn discussion with Michał Piszczek, who pointed out that a validation gate becomes another stale label once its review window expires with nobody owning revalidation; then that the owner field rots the same way, since a personal name goes stale the moment someone leaves and nobody transfers the check; and then that a resolved owner is not a responsive one, because an empty rotation still passes a membership check.

A minimal tracer recording token counts, model, and cost as span attributes on an LLM call — the concrete version of this module’s “log AI-specific fields on the same span as the outcome” principle — paired with prompt redaction applied before anything is recorded:

Span-based tracing with token/cost attributes and pre-storage redaction

from __future__ import annotations
import time
from contextlib import contextmanager
from dataclasses import dataclass, field
from typing import Iterator
REDACTED = "[redacted]"
@dataclass
class Span:
name: str
attributes: dict[str, object] = field(default_factory=dict)
start_time: float = field(default_factory=time.monotonic)
duration_seconds: float | None = None
def set_attribute(self, key: str, value: object) -> None:
self.attributes[key] = value
def finish(self) -> None:
self.duration_seconds = time.monotonic() - self.start_time
class Tracer:
def __init__(self) -> None:
self.finished_spans: list[Span] = []
@contextmanager
def start_span(self, name: str) -> Iterator[Span]:
span = Span(name=name)
try:
yield span
finally:
span.finish()
self.finished_spans.append(span)
def redact(text: str, sensitive_substrings: list[str]) -> str:
redacted = text
for substring in sensitive_substrings:
redacted = redacted.replace(substring, REDACTED)
return redacted
tracer = Tracer()
def call_model(prompt: str, model: str, sensitive_substrings: list[str]) -> str:
with tracer.start_span("llm.call") as span:
span.set_attribute("llm.model", model)
span.set_attribute("llm.prompt_tokens", len(prompt.split()))
span.set_attribute("llm.prompt_redacted", redact(prompt, sensitive_substrings))
response = f"response to: {redact(prompt, sensitive_substrings)[:40]}"
span.set_attribute("llm.completion_tokens", len(response.split()))
span.set_attribute("llm.estimated_cost_usd", 0.00002 * len(prompt.split()))
return response

llm.prompt_redacted is computed and stored on the span instead of the raw prompt — redaction happens before anything is recorded, not as an access-control policy layered on top of unredacted data already sitting in a trace store. Recording token counts and estimated cost directly as span attributes means a single trace query answers “which step of this request was slow and what did it cost,” instead of joining a trace store against a separate cost-tracking system after the fact — the same per-request cost tracking Module 4 introduces, now attached to the exact span that incurred it.

The same instinct applied to the pipeline. pytest exiting 0 is the input to this gate, not its conclusion — the job also has to say how many tests it found, how many actually ran, and what it declined to run:

A CI gate that proves the suite ran, rather than that pytest exited 0

.github/workflows/lab-ci.yml
- name: Tests
run: ./.venv/bin/python -m pytest -q --junitxml=results.xml
# Deliberately a separate, independently visible step. Folded into the one
# above, a single `|| true` would disable the suite and its own audit together.
- name: Prove the suite was discovered and run
run: ./.venv/bin/python scripts/assert_suite_ran.py results.xml --floor 19
scripts/assert_suite_ran.py
"""Fail the build when the suite did not run, even though pytest exited 0."""
from __future__ import annotations
import argparse
import sys
import xml.etree.ElementTree as ET
def main() -> int:
parser = argparse.ArgumentParser()
parser.add_argument("report")
parser.add_argument("--floor", type=int, required=True)
args = parser.parse_args()
# Trusted input: this file was written by pytest in the step above, inside
# the same job. Parsing a report from anywhere else -- an artifact download,
# another repository -- wants defusedxml instead; the stdlib parser is
# documented as vulnerable to entity-expansion attacks on hostile XML.
root = ET.parse(args.report).getroot()
suites = root.findall(".//testsuite") or [root]
discovered = sum(int(s.get("tests", "0")) for s in suites)
skipped = sum(int(s.get("skipped", "0")) for s in suites)
executed = discovered - skipped
# Printed unconditionally: the scope of what was validated is the artifact,
# and it belongs in the log whether or not this run is the one that fails.
print(f"discovered={discovered} executed={executed} skipped={skipped} floor={args.floor}")
if discovered == 0:
print("FAIL: collection found no tests at all", file=sys.stderr)
return 1
if executed < args.floor:
print(f"FAIL: {executed} tests ran, below the floor of {args.floor}", file=sys.stderr)
return 1
if skipped:
# Not a failure on its own. Named anyway, because an unavailable service
# hollows a suite out one skip at a time and each one is individually green.
print(f"NOTE: {skipped} test(s) skipped", file=sys.stderr)
return 0
if __name__ == "__main__":
sys.exit(main())

The floor is the weakest part of this and the part most likely to be copied, so treat it as one invariant among several rather than the answer. What actually carries the gate is the zero-discovery check, which never needs tuning, and the printed line — a reviewer comparing discovered=19 against discovered=4 on the previous run learns something no threshold encodes. In a monorepo, add an assertion that the packages you expected to be inspected appear in the tool’s own output, since “linted successfully” and “linted the thing you changed” are different claims.

Making a pipeline's validation scope observable

  • Zero discovered tests fails the build wherever tests are expected.
  • A count floor exists where the suite is stable enough for one to mean something, and a regression against the previous run is checked where it is not.
  • Every run reports discovered, executed, passed, failed, and skipped — in the log, not only in a summary UI that a future migration will drop.
  • An unexpected skip is treated as a signal to look at, not as a pass.
  • Lint, typecheck, unit, and integration run as independently visible jobs, so one swallowed failure cannot cover for the others.
  • The tools assert that the packages and directories you expected were actually inspected.
  • Coverage is enforced as a gate wherever coverage means anything, rather than merely reported.
  • The jobs that matter are required checks in branch protection — a gate nothing enforces is decoration.
  • No step hides a failure behind || true, continue-on-error, or a pipeline that returns the wrong command’s status.
  • The pipeline is itself tested periodically by introducing a change that must fail, confirming it does, and reverting — the same reasoning as Module 4’s eval harness needing a deliberately broken case scored at zero.
  • Workflow and validation-config changes carry extra review — CODEOWNERS, or a reviewer who did not write the change — because they are the changes that can disable the evidence.

And for the controls whose evidence expires:

  • Approvals and validation results carry a timestamp and the version of what was reviewed.
  • Controls that should not be trusted indefinitely have a stated validity period and an owner expressed as a role, team, or rotation rather than an individual’s name.
  • An owner that no longer resolves expires the control the same way lapsed evidence does — and a resolved owner is not yet a responsive one, since an empty rotation passes a membership check.
  • Controls that cannot be coupled to a path someone must traverse are exercised rather than recorded — an unannounced test, on the channel the control actually uses.
  • A material change to the reviewed artifact invalidates the approval instead of inheriting it.
  • Expiry produces a visible state — expired, requires-review — rather than leaving approved standing.
  • Lapsed high-risk controls fail closed; lower-risk ones warn, so the gate stays worth having.
  • Waivers are dated and justified, because a permanent exception is a control that was deleted without saying so.

An agentic system’s p99 latency SLO starts burning through its error budget. The RED metrics show elevated duration on the top-level request span alone — not enough to diagnose anything. Tracing into the nested llm.call and tool.call spans (this module’s Architecture diagram) shows the slowdown is isolated to one specific tool call, not the model calls at all — a downstream API the tool wraps has degraded. Without span-level tracing through the agent’s tool-call fan-out, this would have looked like a generic model latency regression, sending the on-call engineer to look at the wrong system entirely.

Broken trace context propagation across an async boundary

A background task or queue-dispatched job started without explicitly carrying the current trace context produces an orphan span with no parent — the trace looks incomplete, and there’s no error to point at why, only a gap where a step should have appeared. This is a direct consequence of async code’s implicit-context assumptions, per Module 1.

Unredacted prompts and responses reaching the log or trace store

If redaction happens only as an access-control policy on who can query stored logs, rather than before storage, the sensitive data is already sitting in a system that policy alone can’t retract — any breach or overly broad access grant to that store exposes it. This module’s Deep Dive and Implementation sections both treat redaction as a pre-storage step, not a downstream control.

Cardinality explosion in metric labels

Adding a high-cardinality label to a metric (a raw user ID or request ID as a label, rather than a bounded dimension like tenant_tier) multiplies the number of distinct time series the metrics backend has to store, often silently, until query performance or storage cost degrades enough to notice — usually far after the label was added, making the root cause hard to trace back.

Alerting on symptoms instead of SLO burn rate

Alerting on every individual metric threshold crossing (a single slow request, one error) produces enough noise that on-call engineers start ignoring pages — the specific failure mode SLO burn-rate alerting (this module’s Deep Dive) exists to prevent, by distinguishing a blip from an actual trend toward exhausting the error budget.

Treating a quality regression as invisible because nothing else alerted

A model or prompt change that degrades answer quality produces no RED-metric symptom at all — the requests are fast and return 200s. Without the eval harness from Module 4 running as an independently monitored signal, this regression can ship and stay live indefinitely, precisely because traditional observability has no way to see it.

A green pipeline that validated nothing

Every one of these exits 0. None of them has any runtime symptom, and each is a real way a suite stops running while the badge stays green:

  • Discovery returns nothing. A moved directory, a renamed package, a changed rootdir or testpaths, and collection finds zero tests. Zero passing tests is a pass.
  • The count quietly drops. A path or config change leaves the suite running, but a third of it is no longer collected. Nothing distinguishes this from tests having been deleted on purpose.
  • Cases skip themselves. An absent secret, an unreachable service, a missing fixture or an unset environment variable turns case after case into a skip. Skips are not failures, so a suite can become a no-op one condition at a time.
  • The checker is aimed somewhere else. Typechecking a directory you no longer ship, or linting a path list that never gained the module you added, reports clean on code it never opened.
  • A filter excludes the thing you changed. Monorepo path filters and conditional if: guards on integration or end-to-end jobs are both ways for the affected package to be the one package not checked.
  • Failure is swallowed on purpose. A stray || true, a continue-on-error, or a shell pipeline that returns the exit status of the wrong command converts a failing gate into a passing one, usually added to unblock something and never removed.
  • Coverage collapses without comment, because coverage is reported rather than enforced.
  • The required checks change. Generated or refactored workflow config alters which jobs branch protection actually requires, so a gate stops being a gate while still appearing in the run list.

The compound version, and the one to watch for in review: a single change that edits both an implementation and the mechanism that validates it. Each half looks reasonable; together they can remove the evidence that the other half is wrong.

Full tracing vs. sampled tracing

Tracing every request gives complete visibility but costs proportionally more in storage and instrumentation overhead as traffic grows. Sampling (recording a percentage of requests, or always recording errors and slow requests specifically) keeps cost bounded at high volume, at the risk of missing the specific request an investigation actually needs if it wasn’t sampled.

Structured logging verbosity vs. log volume cost

Logging more fields per request (full prompt/response, every intermediate step) makes debugging easier but multiplies log storage and ingestion cost at scale. The usual resolution is tiered: rich structured logs at a sampled rate, terser logs for every request.

Real-time eval monitoring vs. periodic batch evaluation

Running the eval harness continuously against live traffic catches quality regressions faster but costs compute on every request path or a shadow path. Periodic batch evaluation against a held-out sample is cheaper but detects a regression only as fast as the batch cadence allows — the right choice depends on how costly a live quality regression actually is to leave undetected.

A minimum test count vs. the several invariants that actually hold a pipeline honest

A test-count floor is the cheapest defence against a suite that silently stops running, and it is the one most likely to be adopted as though it were the whole answer. It ages badly: tests get parameterised, consolidated, split, and moved, so a fixed number produces failures that are someone’s refactor rather than someone’s bug — and a floor that cries wolf gets raised until it bounds nothing.

The floor earns its place as one invariant among several. Non-zero discovery never needs tuning and catches the worst case outright. A regression check against the previous run’s count catches the drop a floor set months ago would sit underneath. Asserting that the suites and packages you expected are present catches the filter that excluded them. Coverage or behavioural gates catch the tests that ran but stopped asserting. Skip visibility catches the suite hollowing out one condition at a time. Independent, separately required jobs stop one swallowed failure from covering for the rest. Any single one of these is defeatable; the set is not, and the set is cheap.

  • Redact before storage, not after — this module’s Deep Dive and Implementation sections both treat this as non-negotiable; an access-control policy on a logging system doesn’t undo having stored sensitive data unredacted in the first place.
  • Observability data itself needs access control — traces and logs routinely contain enough detail (request payloads, internal service topology, tenant identifiers) that a database of telemetry is itself a target, not just a debugging convenience with no attack surface of its own.
  • Metric and log exporters crossing a network boundary need the same transport security as any other traffic — telemetry is not exempt from encryption-in-transit requirements just because it’s “just monitoring data.”
  • Instrumentation itself has overhead — every span, every structured log field, every metric emission costs CPU and, for traces, serialization and network cost to the collector. This is usually small per request but adds up at scale, which is exactly why sampling strategies (this module’s Trade-offs) exist rather than treating “trace everything” as a free default.
  • High-cardinality metrics slow down the metrics backend itself — this module’s Failure Modes section covers the failure; the performance angle is that query latency against a cardinality- exploded metric degrades for everyone querying that backend, not just the team that added the problematic label.
  • Trace-based debugging is only as fast as trace query latency — a trace store that takes minutes to return a query result during an active incident defeats the purpose of having traced the request in the first place.
  • Log and trace volume scale with request volume, not linearly with usefulness — the marginal value of the millionth identical successful request’s full trace is close to zero, which is the practical argument for sampling (or always-sample-errors-and-slow-requests) once volume passes a threshold where full retention stops being affordable.
  • Metric cardinality has to be planned, not discovered — a labeling scheme that works at low scale needs deliberate bounds (a fixed, small set of allowed label values) before traffic grows into the range where an unbounded label turns into the cardinality explosion this module’s Failure Modes section covers.
  • Multi-region deployments (per Module 11) need a per-region or federated observability strategy — a single global metrics/trace backend that every region writes to synchronously reintroduces the cross-region latency and availability coupling Module 11 argues against everywhere else in the architecture.

Why can't traditional observability (metrics, logs, traces) detect an LLM quality regression?

All three signals describe execution — did the request succeed, how long did it take, what path did it take through the system — not whether the answer produced was correct. A model or prompt change that degrades answer quality produces a normal-looking trace and a 200 response; only an independently monitored evaluation signal (Module 4) can catch it.

What's the practical difference between an SLI, an SLO, and an error budget?

An SLI is the specific measured metric (e.g., p99 latency); an SLO is the target for it (e.g., under 800ms over 28 days); the error budget is how much SLO violation is tolerable before it becomes unacceptable. Alerting on error-budget burn rate distinguishes a temporary blip from a trend that will actually exhaust the budget, producing fewer, more actionable pages.

How would you trace a slow request through an agent that makes several tool calls?

Propagate trace context through every model call and tool invocation as nested spans under the top-level request span, per this module’s Architecture diagram — this turns “the request was slow” into “this specific tool call was slow,” which a single aggregate latency metric on the outer request can’t distinguish on its own.

How do you log LLM prompts and responses for debugging without creating a data-exposure risk?

Redact sensitive content before it’s written to any log or trace store, not as an access-control policy applied afterward — this module’s Implementation section applies redaction inline, on the same span attribute that also records token counts and cost, so nothing unredacted is ever persisted in the first place.

Your CI has been green for three weeks. What would convince you the tests are actually running?

Nothing about the green itself — a pipeline that skipped its suite and one that passed it emit the same signal, and exit code 0 only says the process it invoked considered itself successful. What convinces me is the run reporting how many tests it discovered and executed, and that number being what I expect for the code that changed. Failing on zero discovery costs nothing and catches the worst case; a count regression against the previous run catches the quieter one. The strongest version is periodically breaking something on purpose and confirming the pipeline goes red, because a check that has never failed is indistinguishable from a check that cannot fail.

SLO-Driven AI Operations implements this module’s SLO material as a running control loop: error-budget tracking with multiwindow burn-rate alerting, an autoscaling controller that treats a fast burn differently from ordinary load, and runbook execution that escalates rather than retrying. It is labelled production-shaped — the control logic is real and tested, the metrics source and scaling target are stand-ins.

Prove redaction happens before storage, not after

Using this module’s Tracer and call_model, write a test that calls call_model with a prompt containing a marker string (e.g., "api_key: sk-test-123") in sensitive_substrings, then assert that no finished span’s attributes contain that marker string anywhere — proving redaction happened before the span was recorded, not merely that a redaction function exists somewhere in the codebase.

Version Date Change
1.0.0 2026-08-07 Initial publication.
1.1.0 2026-09-02 Added false-green CI as an observability problem — the four claims behind a trustworthy green build, the failure classes that produce a false one, a gate that reports discovered/executed/skipped, why a test-count floor is one invariant rather than the answer, and the “validate the validator” contract for AI-assisted changes. Sharpened by feedback from Dushyant K..
1.2.0 2026-09-02 Extended it one layer up — trust has an expiry. A status can be wrong because its evidence went stale, not only because a check did not run: the owner/validity/expiry/revalidation questions, approvals bound to the artifact version they reviewed, and the observation that this handbook’s own freshness linter is already that pattern. From a LinkedIn discussion with Michał Piszczek.
1.3.0 2026-09-02 Closed the gap the previous version left — the owner field rots the same way the label does. Ownership has to point at a role, team, or rotation rather than a person, and an owner that no longer resolves expires the control just as lapsed evidence does. Continuing the LinkedIn discussion with Michał Piszczek.
1.4.0 2026-09-02 Named the regress the previous three versions were walking down and said where it stops — a resolved owner is not a responsive one, since an empty rotation passes a membership check. Adds a ranking of liveness signals, coupling rather than another rung as what terminates the tower, exercise as the only evidence where coupling is unavailable, and liveness as a property of the channel rather than the team. Writing it surfaced the same gap here — main had no branch protection, so none of this repository’s CI checks were required — which is now fixed, admins included. Continuing the LinkedIn discussion with Michał Piszczek.