Agent Identity Broker
Read the transcript
1. Narrowing is three things, not one
Host: So we’re digging into this agent identity broker lab today, and the headline claim is almost provocative: narrowing isn’t one thing, it’s three things, and if you only do one of them, you don’t actually have a control. Walk me through why that split even matters.
Guest: Right, the instinct is to think of a token exchange as ‘shrinking’ access, like there’s a single dial you turn down. But the lab shows audience, scope, and lifetime are independent axes, and there’s a test that asserts all three move together. The reason that matters is the counterfactual: take a token that’s correctly bound to one server, so audience looks perfectly narrowed, but let it still carry every scope the user originally held. That token opens both of that server’s tools — including the refund tool — because audience only decided where the token can go, not what it’s allowed to do once it’s there.
Host: So a design review could look at that audience binding, see it’s tight, and sign off — while the refund tool is sitting wide open behind it. That’s the trap. So what does the broker actually do when it’s asked to hand out scopes the caller was never granted, or handed a token that wasn’t even addressed to it in the first place?
Guest: It refuses, flatly, at exchange time. Ask for a scope the subject doesn’t hold and the request is denied — a narrowing mechanism that can widen isn’t a control, it’s a suggestion. And the harder refusal is the second one: the broker checks that the presented subject token was actually addressed to it before minting anything, because without that check it’s just an escalation service — hand it any stolen token and walk away with a correctly signed credential for whatever target you like.
2. What one token actually opens
Host: So walk me through the actual measurement — you built a fleet, three servers, six tools total, and you ran real tokens against it to see what opens. What’s the headline number for the properly scoped credential?
Guest: One tool, out of six. That’s the agent-scoped credential — one audience, one scope, and it opens exactly the door it was minted for and nothing else. But narrow only the audience and leave scope alone, and the number jumps to two, because now it’s both tools on that server, refund included, since scope never said no.
Host: And the row everyone wants to misread is the raw forwarded user token opening zero — that sounds like a win for passthrough, doesn’t it?
Guest: It sounds that way, and it’s exactly backwards. It opens zero here because its audience names the identity provider, not this fleet — every server here happens to reject it, but hand that same token to whoever’s on the other end of that forward and it’s still a fully live user credential wherever it does match. The zero measures this fleet’s luck, not the token’s safety.
3. A protection that has never failed can’t be trusted
Host: So if that zero was really just luck, how do you actually prove the audience check is doing anything at all, rather than just sitting there looking like a control?
Guest: You break it on purpose. Set verify_aud to False in the resource server’s decode options, the exact thing a tired engineer might flip during a debugging session, and run the suite — six tests fail across all three files. Flip it back and they all pass again, and that flip is the whole point, because audience verification has no runtime symptom when it’s missing. The system keeps running, keeps serving requests, looks completely healthy, it’s just unprotected, so a test suite that doesn’t visibly die when the check disappears was never testing the check in the first place.
Host: So the mutation itself has to be a permanent part of CI, not a one-time exercise you ran once and trusted forever.
Guest: Exactly, it’s its own job in the pipeline — disable the check, assert the suite goes red, and fail the build if it doesn’t. Because a security test that has never once failed is indistinguishable from a test that can’t fail, and that’s the whole thesis: narrowing audience, scope, and lifetime is only a control if something breaks loudly the moment it’s gone, otherwise it’s a design review that passed on vibes.
Not covered
The planner wanted these and found nothing in the source to support them:
- How the broker would integrate with a real, non-simulated identity provider in production
- The principal-level design questions about where the exchange should live (host vs. sidecar) and who decides minimum scope
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Module 15 argues that a short-lived, audience-bound, scope-narrowed credential bounds an agent’s blast radius where a forwarded user token does not. That is a comparative claim, and a comparative claim is worth exactly as much as its measurement. This lab measures it.
Source: labs/agent-identity-broker
The path a token takes
Section titled “The path a token takes”flowchart TD
subgraph Broker["TokenBroker — simulated authorization server"]
Sub["Subject token<br/>aud = issuer · 5 scopes · 8 hours"]
Sub --> Guard{"Presented to the right audience?<br/>Scopes a subset?<br/>Actor listed in may_act?"}
Guard -->|No| Denied["ExchangeDenied"]
Guard -->|Yes| Mint["Agent token<br/>aud = one server · 1 scope · 5 minutes<br/>act = agent"]
end
Mint --> C1
subgraph RS["ResourceServer — five checks, every request"]
C1["1 · signature"] --> C2["2 · issuer"]
C2 --> C3["3 · expiry"]
C3 --> C4["4 · audience names this server"]
C4 --> C5["5 · scope covers this tool"]
end
C5 --> Run["Tool runs<br/>audit: agent acting for user"]
C1 -.-> Rej
C2 -.-> Rej
C3 -.-> Rej
C4 -.-> Rej
C5 -.-> Rej
Rej["TokenRejected<br/>401 / 403 on the wire<br/>reason goes only to the log"]What it demonstrates
Section titled “What it demonstrates”Narrowing is three things, not one. The exchange bounds audience, scope, and lifetime together.
test_exchange_narrows_audience_scope_and_lifetime_together asserts all three, because a system
that narrows only the easiest one still reads as compliant in a design review.
Audience limits where; only scope limits what. The most useful test here is the counterfactual: a token bound to one server but carrying every scope the user has opens both of that server’s tools, including the refund. Half the discipline gets you half the protection, and the measurement says so.
A narrowing mechanism that can widen is not a control. Requesting scopes the subject does not hold is denied at exchange time.
The broker refuses to launder tokens. It verifies the presented subject token was addressed to it before agreeing to mint anything. Without that check the broker is an escalation service: present any stolen token, receive a correctly-signed credential for a target of your choosing.
The measurement
Section titled “The measurement”blast_radius.measure() walks one token against every tool on every server in a three-server,
six-tool fleet and reports what it opens:
| Credential | Tools opened, out of 6 |
|---|---|
| Agent-scoped: one audience, one scope | 1 |
| Audience-narrowed, scope not narrowed | 2 — both tools on that server, refund included |
| The raw user token, forwarded | 0 — no server accepts it |
That last row is the one worth reading carefully, and it is not the flattering result. A forwarded user token opens nothing in this fleet, because its audience names the identity provider and every server rejects it. Passthrough is not dangerous because it over-grants directly; it is dangerous because of what it hands to whoever receives it. The lab measures what it can measure and does not dress the number up.
Prove the tests can fail
Section titled “Prove the tests can fail”The exercise the lab is really for. Disable audience verification the way a debugging session
would — "verify_aud": False in the resource server’s decode options — and six tests fail across
all three files. Restore it and they pass.
That matters because audience verification has no runtime symptom when it goes missing. The system stays fully functional and completely unprotected, which is why a test that fails when the check disappears is the only thing actually holding the protection in place.
CI runs that mutation as its own job: it disables the check, asserts the suite goes red, and fails the build if everything still passes. A security test that has never failed is indistinguishable from one that cannot.
Run it
Section titled “Run it”cd labs/agent-identity-brokeruv venv .venv && uv pip install --python .venv/bin/python -e '.[dev]'./.venv/bin/python -m pytest -qVerify it
Section titled “Verify it”The three gates every lab here is measured against, plus the mutation:
./.venv/bin/python -m ruff check ../.venv/bin/python -m mypy src./.venv/bin/python -m pytest -qTwenty-two tests, ruff clean, mypy --strict clean.
Principal-level discussion points
Section titled “Principal-level discussion points”- Where does the exchange belong — in the agent host, or in a sidecar every agent calls? The host is simpler; the sidecar is the only version that gives you one place to enforce policy and one place to look during an incident.
- What is the minimum scope, and who decides? The lab hard-codes it. In a real system this is the design work that does not go away, and the failure mode is a slow drift back toward “whatever the user has.”
- Five-minute tokens and long-running tasks. Expiry mid-task is a handled path or an outage. The lab’s fourth exercise makes you pick.
- Attribution versus privacy. The
actclaim makes agent actions attributable. That is a security win and a data question, and the second one usually arrives later than it should.
Related
Section titled “Related”- Architecture: Agent Identity Platform — the design review of what this lab implements, including the blast-radius numbers above read as a failure mode.
- Module 15: Agent Identity and Access — the concepts, the specifications, and the failure modes this lab implements.
- Policy-Gated Tool Runtime — the authorization half: what a caller may do, once you know who they are.
- Module 6: MCP — the protocol whose authorization requirements this implements, and why credentials belong on the transport.