Architecture: Agent Identity Platform
Read the transcript
1. The Question Authorization Never Answers
Host: So we’ve spent a lot of time in this space talking about policy checks — what a tool is allowed to do, what scopes it needs, all of that. But there’s a question sitting underneath all of it that I don’t think gets asked enough: when your agent calls a tool, who exactly is calling? Today we’re digging into that, because it turns out the answer everyone reaches for first is the answer that quietly breaks everything.
Guest: Right, and the reason it’s so tempting is that policy-gated execution already assumes an identity — it scopes every permission to a caller, but it never actually manufactures that caller. So the obvious move is to just hand the agent the user’s own token. It’s zero extra work, it runs immediately, and the agent inherits exactly what the user could already do. The problem is that an agent asked to summarize your inbox is now holding the same credential that can issue a refund or delete a record, because that token was never scoped to the task — it was scoped to the person.
Host: And that one shortcut breaks three separate things at once, from what I understand — blast radius, attribution, and revocation — and it breaks them silently, which is the part that should worry people. So that’s where we’re headed this episode: what it actually takes to mint an identity for the agent itself, one a resource server can distrust by default, so that fixing one of those three doesn’t mean sacrificing the other two.
2. What a Real Fix Requires
Host: Okay, so if forwarding the user’s token is the shortcut that quietly breaks everything, what does the fix actually demand? You mentioned mint an identity for the agent — what does that checklist look like in practice?
Guest: It comes down to seven things that all have to hold at once. First, agent-scoped credentials — every call carries a token minted for that agent and that task, never the user’s token relayed onward. Second, narrowing has to happen in three dimensions together, audience, scope, and lifetime, because narrowing just one isn’t narrowing at all. Third, the resource server has to do its own validation every time — signature, issuer, expiry, audience, scope — regardless of what the caller claims about itself. Fourth, the exchange that produces the agent’s token has to be incapable of widening; ask for a scope the subject doesn’t hold and the broker refuses. Fifth, that broker can’t launder tokens either — it has to verify the presented subject token was actually addressed to it before minting anything downstream. Sixth, the resulting credential has to name both the agent and the user it’s acting for, so the audit trail can tell them apart. And seventh, revocation has to work short of the user — stopping one agent must not mean disabling the human it’s acting for.
Host: That’s a longer list than I expected for what sounded like one shortcut. Which of these seven is the one people actually get wrong in practice — the one the rest of this episode is going to keep circling back to?
3. The Constraints That Don’t Show Up Until They Bite
Guest: It’s audience verification, hands down. Everything else on that list breaks loudly — a broker that’s down fails every tool call. But skip audience verification and nothing breaks. The system runs perfectly, every demo passes, every test goes green, because the only thing missing is the check that would have said ‘this token wasn’t issued for you.’
Host: So there’s no dashboard, no error, no degraded mode — it just silently works right up until it’s exploited. Walk me through why a resource server would even skip that check. It sounds like the one thing you’d never forget.
Guest: Because the client hands you a token and says ‘this is scoped to me,’ and it’s tempting to just believe that self-report — but a claim the caller makes about itself is worth nothing, it’s a suggestion box, not authorization. The actual spec, MCP’s 2026-07-28 authorization update, requires the server to independently validate audience, not trust the assertion. And it gets worse when you add short-lived tokens into the mix, because a five-minute lifetime is great security and a real hazard for a task that runs an hour — expiry mid-task is either a path you’ve built for or an outage you didn’t see coming.
4. Two Stages, Two Jobs: Narrow, Then Distrust
Host: So walk me through what actually happens end to end, because you keep saying two stages and I want to know why that split matters so much.
Guest: The broker’s whole job is narrowing — it takes a long-lived, broadly-scoped subject token and mints something tight: one audience, one scope, a few minutes of life. But before it mints anything it checks three things — was this subject token actually addressed to the broker itself, are the requested scopes a subset of what the subject holds, and is this actor even permitted to act for this subject. Only after all three pass does it hand over the narrow credential, and that’s the last the broker ever hears about it.
Host: And the resource server doesn’t just take the broker’s word that the narrowing was done right.
Guest: Right, it distrusts by design — it re-derives signature, issuer, expiry, audience, scope, all of it, from the token in front of it, and the only thing it ever got from the broker’s world is the issuer’s public key. It never phones home to ask ‘is this good,’ because a check that depends on a live call to the authority is a check that fails open the moment that authority is slow or down. And when something fails, the caller gets a bare 401 or 403 — the actual reason, which of the five checks blew up, goes only to the log, because handing that back to the client is a free oracle for figuring out exactly how to narrow the attack.
5. How This Fails: Half-Narrowed, Laundered, and Drifted
Host: So let’s make this concrete. If passthrough is the disease, what does the actual damage look like when someone finally measures it?
Guest: The passthrough case is the worst because the credential is good everywhere the user is — it’s not scoped to one server, it’s the user’s whole reach, so whoever catches it inherits the user’s blast radius. But even the half-measure fails in a way that looks fine on paper. In the lab, a token narrowed to a single server but still carrying every scope the user holds opens two of six tools on that fleet — both tools on that one server, including the refund tool. You’d sign off on that in a design review because ‘we scope tokens per service’ is technically true, and it’s still handing out the one tool you most needed to lock down.
Host: And the broker itself — that’s supposed to be the trusted narrowing point. What happens when it skips its own check?
Guest: If the broker doesn’t verify the audience on what’s handed to it, it stops being an identity service and becomes an escalation service — feed it any stolen token from anywhere, get back a correctly-signed credential for whatever target you pick. And the quieter version of that is someone flipping the verify-audience setting to false to unblock a local error and never flipping it back. Every test keeps passing, every request keeps succeeding, there’s no symptom at all — the only thing that catches it is a test built to fail the moment that check is gone. Scope creep works the same way on a longer clock: minimum scope gets widened every time a task fails for want of a permission, and it never narrows back because nothing breaks when it’s too wide, so eighteen months in the agent’s token is just the user’s token with extra steps.
6. The Broker’s Scaling Problem
Host: So we’ve established short lifetimes are the security lever, but there’s a bill attached to that lever, right? Walk me through what actually happens to the broker when you turn it.
Guest: Halve the token lifetime and you roughly double the exchange volume hitting the broker, because every agent action now needs a fresher credential more often. That’s the central tension of this whole design — a security dial that’s wired directly to a capacity number. If you don’t think about it as a capacity number, you’ll tune security and take down the platform by accident. The thing that saves you is that exchanged tokens are cacheable, keyed by subject, actor, audience, and scope, and they get evicted before expiry rather than left to expire naturally. That decouples broker load from raw tool-call volume — but it also means that cache is now a credential store, with everything that implies about how you guard it.
Host: And resource servers don’t share that problem the same way — they’re not calling home on every request?
Guest: Right, they scale independently because validation is entirely local — signature, issuer, expiry, audience, scope, all come from the token itself plus a cached public key. No coordination, no shared state, no call back to the broker. The one place they still depend on the outside world is JWKS — cache the keys, refresh on an unknown kid, and rate-limit that refresh, or a key rotation turns your whole fleet into a thundering herd against the identity provider. But none of that changes the fact that the broker itself is a single point of failure for every agent action — it needs the availability budget of the platform’s most critical path, because functionally, that’s exactly what it is.
7. The Security Checklist and Proving the Check Can Fail
Host: Let’s make the checklist concrete instead of decorative — what are the rules that actually carry weight, the ones where skipping them doesn’t show up as an error, it shows up as a breach nobody notices?
Guest: Validate audience server-side on every request, and verify it at the broker before minting anything, so the broker itself can’t be used to launder a credential into a wider one. Refuse to widen — requested scopes have to be a subset of held scopes, denied at exchange time, not caught later at use time — and bind audience, scope, and lifetime together, because a test that only asserts the easy dimension narrowed will happily pass on a system that only narrows the easy dimension. Never return the rejection reason to the client, since five checks with five distinct error messages is a map you’re handing an attacker for free.
Host: And audience is the one you keep coming back to as the load-bearing one — presumably because if it’s silently missing, nothing breaks, nothing alerts, the system just runs fully functional and fully unprotected.
Guest: Exactly, which is why the lab doesn’t just test it once and trust the test — it proves the test can fail. Turn audience verification off in the resource server’s decode options, the way a tired engineer debugging something unrelated actually might, and six tests go red across all three files; restore it and they pass again. CI runs that exact mutation as its own job — disable the check, assert the suite goes red, fail the build if it doesn’t — because a security test that has never failed is indistinguishable from one that can’t.
8. The Trade-offs: Lifetime, Placement, Granularity, Privacy
Host: Let’s talk trade-offs, because none of this comes free. Start with lifetime — why not just pick one duration and move on?
Guest: Because a single global lifetime is always wrong in two directions at once. Five minutes is defensible for a refund scope where a leaked credential could do real damage, but an hour is fine for a read-only search — deciding per scope is what keeps that dial honest instead of splitting the difference badly for everything.
Host: And the same tension shows up in placement, audience granularity, and even attribution itself, doesn’t it?
9. What All This Actually Costs
Host: So let’s put a price tag on all of this. When people first hear ‘exchange a token on every tool call,’ their instinct is that it’s expensive — an extra network round trip sitting right on the critical path of every action. Is that the real cost, or is that a red herring?
Guest: It’s real but it’s not the big one, and caching is what tames it — you exchange once, reuse within the token’s lifetime, so short lifetimes stay affordable instead of becoming a tax you pay on every call. The actual expense is broker availability, as we touched on earlier. Meanwhile validation is nearly free — checking a signature against a cached key is microseconds, so ‘we validate on every request’ sounds scary and isn’t. The number that actually surprises people is audit volume, because it scales with agent calls, not user requests, and one request fanning out into dozens of calls means your log volume is easily an order of magnitude bigger than anyone budgeted for.
10. What to Watch on a Dashboard
Host: So audit volume tells you the size of your logs, but not what to actually watch for. If someone’s staring at a dashboard, what are the five or six numbers that tell them apart between an incident, a bug, and a Tuesday?
Guest: Start with exchange denials, broken out by reason, not lumped together. Wrong audience on the subject token is a possible laundering attempt, scope not a subset is a mis-scoped agent, actor not permitted is a delegation misconfiguration — three totally different pages to call, and if you aggregate them into one denial count you can’t tell which one fired. Same logic on the resource server side: which of the five checks failed matters enormously, because an expiry rejection is just operational noise while an audience rejection is either a bug or an attack, and if rejections cluster right around expiry that’s almost always tasks outliving their tokens, not anything malicious. Then there’s scope breadth per agent, which is the one you have to watch as a trend rather than an alert since drift creeps in slowly, and exchange rate per agent, where a sudden spike means either a broken cache or an agent stuck in a loop — either way you want to catch that yourself before the identity provider’s rate limiter does it for you.
11. The Measured Fleet: What a Token Actually Opens
Host: So let’s put actual numbers on this. You’ve got a six-tool fleet across three servers, and you ran three different credentials through it. What did a fully narrowed, agent-scoped token open?
Guest: One tool. Out of six. That’s the token with a single audience and a single scope, exactly what the agent needed for the task it was doing — nothing else on that server or any other server even parses it correctly, let alone honors it.
Host: And when you only narrow the audience, leaving scope untouched — that’s the half-measure you warned about earlier?
Guest: Right, and it opens two, both tools on that one server, refund included. That’s the whole point of the counterfactual — audience tells the token where it’s allowed to go, but only scope tells it what it’s allowed to do once it’s there. Now here’s the row that trips people up: forward the raw user token unmodified, and it opens zero. Every server rejects it outright, because its audience names the identity provider, not them. It looks like the safest row on the table, and it’s actually the most dangerous, because that zero isn’t protection — it’s just the wrong audience field. Passthrough isn’t dangerous because it over-grants directly here; it’s dangerous because of what it hands to whoever receives it, whether that’s a sloppy server or a broker willing to launder it.
12. Defending the Design Under Questioning
Host: Let’s do the lightning round before we close, because these are the questions that actually separate someone who’s absorbed this from someone who’s just nodded along. Why is forwarding the raw token dangerous beyond just over-granting at one server, and what’s the real difference between what audience limits and what scope limits?
Guest: Passthrough makes the agent’s permissions identical to the user’s, so every action gets logged as a human action, and revocation has no lever short of disabling the person — the danger isn’t one over-permissive server, it’s that the credential is good everywhere the user is, so whoever receives it holds the user.
Host: So the whole argument comes back to one sentence: a policy check is only as honest as the caller it’s checking, and everything we’ve walked through — narrowing, distrust, the audience check, the CI job that proves the check can fail — exists to make that caller true instead of assumed. That’s the episode. Thanks for building this out with us.
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Problem
Section titled “Problem”Policy-Gated Tool Execution answers what a caller may do. It assumes something it never supplies: a caller. Every control in that architecture is scoped by identity, which means the whole enforcement pipeline rests on a question it does not itself answer — who is this?
The default answer is to forward the user’s token. It is the easiest thing to build, it works immediately, and it makes the agent’s effective permissions identical to the user’s. An agent asked to summarise an inbox holds the credential that can also issue refunds, because that credential was never about the inbox — it was about the person.
Three things break at once, and they break quietly:
- Blast radius. A compromised or merely over-persuaded agent acts with the user’s full permission set, not with the permissions the task needed.
- Attribution. The audit log says the user did it. Every action an agent takes is recorded as a human action, which is exactly wrong at the moment anyone needs the log.
- Revocation. The only lever is the user’s own credential. Stopping the agent means locking out the person.
What is needed is a credential minted for the agent, narrowed to the task, validated by the server that receives it — and an architecture in which none of those three is optional.
Requirements
Section titled “Requirements”- Agent-scoped credentials. Every agent call carries a token minted for that agent and that task, not the user’s token relayed onward.
- Narrowing in three dimensions together — audience, scope, and lifetime. Narrowing one is not narrowing.
- Server-side validation. Each resource server verifies signature, issuer, expiry, audience, and scope itself, on every request, regardless of what the caller claims.
- An exchange that cannot widen. Requesting a scope the subject does not hold is refused at the broker.
- A broker that will not launder. The broker verifies the presented subject token was addressed to it before minting anything.
- Attribution across the hop. The resulting credential names both the agent and the user it acts for, so the audit record distinguishes them.
- Revocation short of the user. Stopping one agent must not require disabling the human.
Constraints
Section titled “Constraints”- The client’s claim about itself is worth nothing. A resource server that trusts the caller’s
assertion of audience or scope has implemented no authorization at all — it has implemented a
suggestion box. MCP’s
2026-07-28authorization specification requires the server to validate the audience; this architecture treats that as the load-bearing requirement it is. - Short lifetimes and long tasks are in direct tension. A five-minute token is a good security property and an operational hazard for a task that runs for an hour. Expiry mid-task is either a handled path or an outage; there is no third option.
- Audience verification has no runtime symptom. Remove it and the system stays completely functional and completely unprotected. Nothing degrades, no error surfaces, and no dashboard moves.
- The broker is on the hot path. It sits between an agent deciding to act and the action, which makes its availability and latency a property of every tool call in the platform.
- Narrowing depends on someone deciding what “narrow” means. The minimum scope for a task is design work that does not go away, and its failure mode is a slow drift back toward whatever the user has.
Request Flow
Section titled “Request Flow”flowchart TD
subgraph Broker["TokenBroker — simulated authorization server"]
Sub["Subject token<br/>aud = issuer · 5 scopes · 8 hours"]
Sub --> Guard{"Presented to the right audience?<br/>Scopes a subset?<br/>Actor listed in may_act?"}
Guard -->|No| Denied["ExchangeDenied"]
Guard -->|Yes| Mint["Agent token<br/>aud = one server · 1 scope · 5 minutes<br/>act = agent"]
end
Mint --> C1
subgraph RS["ResourceServer — five checks, every request"]
C1["1 · signature"] --> C2["2 · issuer"]
C2 --> C3["3 · expiry"]
C3 --> C4["4 · audience names this server"]
C4 --> C5["5 · scope covers this tool"]
end
C5 --> Run["Tool runs<br/>audit: agent acting for user"]
C1 -.-> Rej
C2 -.-> Rej
C3 -.-> Rej
C4 -.-> Rej
C5 -.-> Rej
Rej["TokenRejected<br/>401 / 403 on the wire<br/>reason goes only to the log"]The shape is two stages with different jobs, and conflating them is the common mistake.
The broker narrows. It takes a long-lived, broadly-scoped subject token and mints a credential bound to one audience, one scope, and a few minutes. Before it does, it checks three things: that the subject token was addressed to the broker itself, that the requested scopes are a subset of what the subject holds, and that this actor is permitted to act for this subject.
The resource server distrusts. It re-derives everything from the token in front of it — signature, issuer, expiry, audience, scope — and shares nothing with the broker but the issuer’s public key. It does not ask the broker whether a token is good. A validation that requires a network call to the authority is a validation that fails open the first time the authority is unreachable.
Rejections return 401 or 403 and nothing else. The reason goes to the log, not to the client:
telling a caller which of the five checks it failed is a free oracle for narrowing an attack.
Failure Modes
Section titled “Failure Modes”Forwarding the user's token
The agent’s blast radius becomes the user’s permission set, its actions are recorded as the user’s, and the only way to stop it is to stop the person. The danger is not that passthrough over-grants at one server — it is what it hands to whoever receives it, which is a credential good everywhere the user is.
Narrowing the audience but not the scope
The measured case, and the most instructive one. A token bound to a single server but carrying every scope the user holds opens two of six tools in the lab’s fleet — both tools on that server, including the refund. Half the discipline buys half the protection, and it reads as compliant in a design review because “we scope tokens per service” is true.
A broker that does not check who the subject token was for
Skip the audience check on the incoming token and the broker stops being an identity service and becomes an escalation service: present any stolen token from anywhere, receive a correctly-signed credential for a target of your choosing. The signature on what comes out is perfect. That is the problem.
Audience verification disabled during debugging
Someone sets verify_aud to false to get past a local error and does not set it back. Every test
still passes, every request still succeeds, and every server in the fleet now accepts tokens minted
for any other server. There is no symptom. The only thing that catches this is a test that fails
when the check disappears.
Token expiry mid-task
A five-minute credential and a forty-minute agent run meet in the middle. Handled, this is a re-exchange the agent barely notices. Unhandled, it is a task that dies most of the way through, usually having already performed some of its side effects.
Scope drift
The minimum scope is set once, then widened whenever a task fails for want of a permission. Nothing ever narrows it back, because nothing breaks when it is too wide. Given eighteen months, the agent token is the user token with extra steps.
Scaling
Section titled “Scaling”- Short lifetimes multiply exchange volume. Halving token lifetime roughly doubles broker QPS. This is the central scaling tension of the design, and it is a security dial wired directly to a capacity number.
- Exchanged tokens are cacheable, keyed by
(subject, actor, audience, scope)and evicted before expiry rather than on it. This is what keeps lifetime short without making the broker’s load proportional to tool calls — but the cache is now a credential store, with everything that implies. - Resource servers scale independently because validation is local. Signature, issuer, expiry, audience, and scope all come from the token and a cached public key. No coordination, no shared state, no call back to the broker.
- JWKS caching is the one place a resource server depends on the outside world. Cache the keys,
refresh on unknown
kid, and rate-limit that refresh — otherwise a key rotation turns every server in the fleet into a thundering herd against the identity provider. - The broker is a single point of failure for every agent action. It needs the availability budget of the platform’s most critical path, because that is what it is on.
Security
Section titled “Security”- Validate audience server-side on every request, and treat the check as load-bearing rather than defensive.
- Verify the subject token’s audience at the broker before minting anything, so the broker cannot be used to launder credentials.
- Refuse to widen. Scopes requested must be a subset of scopes held; deny at exchange time, not at use time.
- Bind audience, scope, and lifetime together, and test that all three narrowed — a test asserting only the easy one passes on a system that narrows only the easy one.
- Record the actor and the subject separately, so the audit answers “which agent, acting for whom” rather than “the user did it”.
- Never return the rejection reason to the client. Five checks and five distinct error messages is a map.
- Keep lifetimes short enough that expiry is the primary revocation mechanism, and treat any design needing instant revocation as needing a real revocation path, not a shorter timeout.
A security test that has never failed is indistinguishable from one that cannot
Audience verification is the highest-value check here and the one with no runtime symptom when it goes missing. The companion lab runs a mutation job in CI: it disables the check, asserts the suite goes red, and fails the build if everything still passes. Any platform relying on server-side audience validation should run the equivalent — otherwise the protection rests on a test whose ability to detect its absence has never been demonstrated.
Trade-offs
Section titled “Trade-offs”Token lifetime: security window vs. broker load
A shorter lifetime bounds the damage from a leaked credential and increases exchange volume proportionally. Five minutes is defensible for high-consequence scopes; an hour is defensible for read-only ones. Deciding per scope rather than globally is what keeps the dial honest, because a single global lifetime is always simultaneously too long for the refund and too short for the search.
Broker in the agent host vs. a sidecar every agent calls
In-host is one fewer network hop and one fewer service to run. A sidecar is the only version that gives you one place to enforce policy and one place to look during an incident — and the only one where a fleet of agents written by different teams cannot each invent their own narrowing rules. The cost is a hop on every tool call and a dependency with the availability profile of the platform.
Audience per server vs. audience per tool
Per-server audiences are what OAuth’s resource indicators naturally express, and they are what most identity providers make easy. They also mean the measured failure above — a token good for every tool on the server it names. Per-tool audiences close that gap and multiply the number of distinct credentials an agent holds during a single task. Scope, narrowed properly, is the cheaper answer to the same problem.
Attribution vs. privacy
Naming the agent and the human it acts for makes every action attributable, which is the security win. It also produces a detailed record of what a specific person’s agent did all day, in a system whose retention policy was probably written for service logs. The security argument arrives first and the data question arrives later, which is the wrong order.
- An exchange per tool call is the naive cost, and it is a network round trip on the critical path of every action. Caching exchanged tokens within their lifetime is what makes short lifetimes affordable.
- Broker availability is the real expense. It sits in front of every agent action, so it carries the platform’s highest availability requirement and the redundancy bill that comes with it.
- Validation is nearly free. Signature verification against a cached key is microseconds, which is worth stating because “we validate on every request” sounds expensive and is not.
- Audit volume grows with agent actions, not user actions — a ratio that is easy to underestimate by an order of magnitude, since one user request can become dozens of agent calls.
Observability
Section titled “Observability”- Exchange denials by reason — wrong audience on the subject token, scope not a subset, actor not permitted. These are three different incidents: a possible laundering attempt, a mis-scoped agent, and a delegation misconfiguration.
- Resource-server rejections by which of the five checks failed. An expiry rejection is operational; an audience rejection is either a bug or an attack, and aggregating them into one error rate destroys the distinction.
- Rejections clustered near expiry, which is the signature of tasks outliving their tokens rather than anything malicious.
- Scope breadth per agent over time. The metric that catches scope drift, and the only one here that has to be read as a trend rather than an alert.
- Exchange rate per agent, where a sudden rise means either a cache that stopped working or an agent in a loop — both worth knowing before the identity provider tells you.
Production Deployment
Section titled “Production Deployment”Before real traffic
- No path forwards a user token to a resource server.
- The broker verifies the subject token’s audience names the broker.
- Requested scopes are checked as a subset of held scopes, and widening is denied at exchange time.
- Audience, scope, and lifetime are narrowed together, with a test asserting all three.
- Every resource server validates signature, issuer, expiry, audience, and scope locally.
- A test fails when audience verification is disabled, and CI runs that mutation rather than trusting that it would.
- Rejection reasons go to the log and never to the client.
- Token lifetime is set per scope, and mid-task expiry has a handled re-exchange path.
- JWKS is cached, refreshed on unknown key ID, and that refresh is rate-limited.
- The audit record names the agent and the subject separately.
Hands-on Lab
A running implementation: token exchange that narrows audience, scope, and lifetime together,
five-check server-side validation, and a blast_radius measurement across a three-server,
six-tool fleet — one tool opened by a fully narrowed token, two when only the audience is narrowed.
CI runs a mutation job that disables audience verification and fails the build if the suite stays
green. The identity provider is simulated, and the README lists what that omits individually
rather than in summary.
Read the lab documentation →
labs/agent-identity-brokerproduction-shaped
Interview Questions
Section titled “Interview Questions”An agent needs to call a tool on the user's behalf. Why not forward the user's token?
Because it makes the agent’s permissions identical to the user’s, records every agent action as a human action, and leaves revocation with no lever short of disabling the person. The specific danger is not over-granting at one server — it is that a forwarded credential is good everywhere the user is, so whoever receives it holds the user.
You narrow the token's audience to one server. What have you actually bounded?
Where the token works, and nothing about what it does there. A token bound to one server but carrying the user’s full scope set opens every tool on that server — in the companion lab’s fleet, two of six, including a refund. Audience limits where; only scope limits what. A candidate who treats per-service tokens as the finished answer has stopped at the halfway point.
What stops a token broker from becoming an escalation service?
Verifying that the subject token presented to it was addressed to it. Without that check, anyone holding any token from anywhere can exchange it for a correctly-signed credential aimed at a target of their choosing — and the output is cryptographically perfect, so nothing downstream can tell. The broker’s own audience check is the thing standing between an identity service and a laundering service.
Someone disables audience verification while debugging and forgets to restore it. How do you find out?
Not from production, because there is no symptom: every request still succeeds and the system is fully functional and completely unprotected. The only reliable detection is a test that fails when the check is absent — and, because a security test that has never failed is indistinguishable from one that cannot, a CI job that disables the check on purpose and fails the build if the suite stays green.
Five-minute tokens, forty-minute agent tasks. Resolve it.
Re-exchange mid-task, and design for it explicitly. The agent holds the subject credential and mints a fresh short-lived token when the current one nears expiry, which keeps the security window small without tying it to task duration. The wrong answers are lengthening the token to cover the worst case — which sets the window by the longest task in the system — and leaving it unhandled, which kills tasks partway through, typically after some side effects have landed.
Where does the broker belong, and what does that choice cost?
In a sidecar every agent calls, if you want one place to enforce narrowing policy and one place to look during an incident; in the agent host, if you want fewer hops and fewer services. The sidecar’s real cost is that it lands on the critical path of every tool call, which means it inherits the platform’s highest availability requirement. The in-host version’s real cost is that narrowing policy is now implemented once per team.