Architecture: Policy-Gated Tool Execution
Read the transcript
1. The core mistake: wiring tool calls straight to dispatch
Host: So let’s start with the mistake that shows up in almost every early agent platform: you take the model’s tool call, and you wire it straight into a dispatch table. The model decides to issue a refund and, well, the system just does it. What could go wrong?
Guest: Everything, because that architecture quietly redefines what your permission system even means. A tool call from a model is a request, not an authorization — but if dispatch treats them as the same thing, then the actual boundary of what your agent can do is no longer defined by your policies. It’s defined by whatever the model can be talked into, by a user, by some retrieved webpage, or even by a tool description someone else wrote.
Host: So the blast radius isn’t ‘the actions we designed for’ — it’s every capability any agent was ever given, full stop. That’s the framing for this whole episode, then: we need something sitting between ‘the model decided’ and ‘the action happened,’ and that something has to answer five separate questions, in a fixed order, instead of collapsing them into one trust decision.
2. Five requirements, five failure modes
Host: Okay, so let’s put names to these five questions. Where do you even start — is it just ‘is this tool allowed,’ or does it get messier than that?
Guest: It starts there but doesn’t end there, and that’s the trap — people think capability scope alone solves it. First question is capability scope: can this caller invoke this tool at all, full stop, not ‘did the model decide to.’ Second is argument validation: every call is checked against the tool’s own schema, server-side, regardless of what the model was shown — no exceptions for what looked plausible upstream. Third is per-tool quota — and the key word is per-tool, because quotas are owned by the tool itself, not shared across the platform, so each tool’s budget stays its own.
Host: And the last two — approval and audit — those feel more like process than architecture. Why do they belong in the same enforcement layer instead of being handled downstream?
Guest: Because if they’re downstream they’re optional, and optional means someone skips them under load. Fourth is human approval for high-risk actions with a bounded wait and an explicit timeout behavior — you decide in advance whether a timeout means deny or escalate, not whichever engineer is on call that day. Fifth is complete audit, including refusals especially, and sixth — because identity has to underlie all five — every decision gets tied to a verified caller, so ‘the agent did it’ is never an acceptable forensic answer.
3. The model is not a trust boundary
Host: You said the model gets tied to a verified identity, but let’s back up — people assume that if you give the model a strict schema and a well-written tool description, it just won’t call things wrong. Why isn’t that a trust boundary?
Guest: Because the schema is advisory, not enforced — it shapes the shape of the model’s output, but nothing stops a compromised or jailbroken model from emitting a call that violates it entirely. And it cuts both ways: the tool description itself is untrusted input, because it flows into the model’s context just like a user message does, so a hostile or compromised tool registry can inject instructions through what looks like harmless metadata. You have to treat every call as coming from an adversarial caller, not just a confused one, or you’ve built your safety net out of suggestions.
Host: So even the caller’s identity in that call can’t just be a field the model fills in and you trust.
Guest: Right, identity has to be verified independently of anything the model asserts, because every single control we’ve described — rate limits, approval, audit — is scoped by who’s calling. If that identity is forgeable, an attacker doesn’t need to break five controls, they just spoof the one thing all five are keyed on and everything downstream silently stops applying to them.
4. Why the check order is load-bearing
Host: So we’ve got scope, validation, quota, approval as a sequence. Is that just a natural pipeline order, or does the order itself matter for security?
Guest: It’s completely deliberate. Scope goes first because it’s the cheapest and most absolute check — if you don’t have a grant on a tool, you shouldn’t even reach argument validation, partly to save work, but also because a schema error message can leak the existence and shape of a tool you had no business knowing about. Validation comes before quota for a similar reason: if you let malformed calls hit the rate limiter first, an attacker can burn through a legitimate caller’s budget with pure garbage, no valid arguments required. And approval sits dead last because it’s the only stage that blocks on a human — you don’t want to pay that latency cost for a call that a cheap check would’ve rejected anyway.
Host: So the ordering isn’t just efficiency, it’s actually preventing specific attacks — disclosure through error messages, budget exhaustion through malformed input.
Guest: Exactly, and the lab frames it the same way — validate-then-authorize and authorize-then-validate are genuinely different systems, not stylistic variants. They differ in what information leaks and whose budget gets spent on a bad request. If you can only describe the five controls as an unordered checklist, you haven’t actually understood the design — being able to defend the specific sequence is the real test.
5. What actually breaks when a check is skipped
Host: Let’s make this concrete. Walk me through what actually happens when one of these five checks just isn’t there — starting with validation, since that’s the one people skip most casually.
Guest: The schema tells the model what a well-formed call looks like, but it guarantees nothing about what actually shows up at the server. Skip validation because ‘the model has the schema’ and you’re one hallucinated argument, or one adversarial client, away from executing something you never meant to accept. And it gets worse — tool descriptions are context too, so a hostile tool source can write a description that steers the model’s behavior far past what the tool does. That’s prompt injection arriving through metadata, and it walks straight past every filter you built for the user’s message.
Host: And the ones further down the chain — approval, quota, audit, identity — those fail just as concretely, not just theoretically?
Guest: Completely concretely. An approval queue with no bounded wait becomes an outage the first time an operator goes unavailable — requests pile up holding connections, and it looks like your backend died, not like policy stalled. Share one quota between a cheap read and an irreversible write, and read traffic can drain the allowance protecting the dangerous operation, or you raise the limit for reads and silently raise it for refunds too. Log only successes and you can’t show an agent kept trying a capability it was never granted — the clearest signal of a compromised caller you’ll ever get, just gone. And none of it matters anyway if the caller identity behind all this is a header any client can set, because then scope, quota, routing, and attribution are all decorative at once, and everything still passes every test you didn’t design to forge it.
6. The trade-offs that don’t have a single right answer
Host: So we’ve established the checks and the ordering — but you keep saying some of these decisions don’t have a single right answer. Give me the first one: fail open or fail closed when the policy store itself goes down?
Guest: It has to be decided per tool, not as a global stance. Fail closed on an irreversible write means a policy-store outage becomes a full outage for that tool, which sounds bad until you compare it to the alternative — fail open silently strips authorization from a refund or a delete, and nobody notices until the damage is done. For a read-only lookup, failing open with loud alerting is legitimate; you’re trading a small window of unchecked reads for continuity, and you can afford that because nothing irreversible is on the table. Same logic applies to the timeout question — deny on timeout and an unavailable approver just produces a visible user-facing failure, which is safe; escalate instead and you’re only justified if the escalation target is actually more reachable, otherwise you’ve just relabeled the hang. The one answer that’s never legitimate is no timeout at all, because that turns a policy decision into an unbounded wait indistinguishable from a hung backend.
Host: And the grant model — per-tool grants versus roles — is that the same kind of judgment call, or is one of those just better?
Guest: It’s a real trade, not a solved problem. Per-tool grants are precise and they audit cleanly — you can point at exactly why a caller could do exactly that action — but they don’t scale once you’ve got dozens of tools, the matrix becomes unmanageable. Roles compress that, but they bring back the classic failure where a role quietly accumulates capabilities nobody actually intended it to have. The way out is roles built from explicit tool grants rather than roles as opaque labels, so you keep the auditability while bounding the growth — and the same deliberateness applies to where you enforce any of this, in the tool runtime so it holds regardless of whether the call comes over MCP, HTTP, or in-process, versus a protocol gateway that’s easier to bolt in front of servers you don’t control but gets bypassed by any path that doesn’t cross it.
7. Making it scale: shared state, stateful approvals, cache invalidation
Host: Let’s talk about running this at scale, because a policy layer that only works on one box isn’t really a policy layer. Where does that first break down?
Guest: Quota. If you implement rate limiting as an in-process token bucket, every replica has its own bucket, so a caller’s real limit is your configured limit times however many replicas happen to be running. That’s not a subtle bug, it’s the limit meaning nothing. The fix in async-ai-gateway is a Redis Lua script that does refill-and-consume as one atomic operation, so all replicas are checking against the same number and the limit actually means what you configured.
Host: And I’d guess approvals have a similar shared-state problem, plus a time dimension.
Guest: Right, approval is a stateful island sitting inside a stateless service, so if the record only lives in the replica that accepted the request, a deploy or a restart just silently drops every human still waiting on a decision — the request doesn’t fail loudly, it just vanishes. Audit has its own version: write volume scales with attempts, not successes, so a refusal storm is precisely the moment logging load spikes, and that’s exactly when you can’t afford the audit path to be the thing that falls over. Then there’s the registry — reads are hot and rarely change so caching is the right call, but that means cache invalidation is the actual mechanism by which a revoked grant takes effect, not the revocation itself.
8. Security posture, cost, and what to watch in production
Host: So if I pull all of this together into a security posture, what’s actually on the list? Not the failure modes anymore, just the concrete things a reviewer should check are present.
Guest: Verified identity from a signed token or mTLS, never a bare header. Server-side validation against the server’s own schema, tool descriptions treated as untrusted with provenance pinned, grants scoped to least privilege with revocation that doesn’t need a redeploy, and an append-only, separately access-controlled audit log that records what was refused, not just that something was refused. Each one closes a specific hole we already walked through — this is just the checklist form of it.
Host: And cost-wise, where does the money and the latency actually go once all that’s running?
Guest: Approval latency dominates because it’s human-scale, so anything you route through it needs to be genuinely consequential or you’re training people to rubber-stamp. Audit storage is the sneaky one — it grows with attempts, not successes, so an attack that generates a thousand refusals is a thousand log writes, which is exactly the wrong moment to be throttled by your own logging bill. On observability, the fix is splitting refusals by reason per caller per tool — scope denial, schema failure, quota exhaustion, approval denial, approval timeout — because lumped into one error rate they’re indistinguishable, and watching approval queue wait-time against the timeout, since a p95 creeping toward it means you’re about to start denying legitimate work.
Host: So the production checklist really is the synthesis of everything we’ve covered — identity, validation, shared rate limits, surviving restarts, tested revocation, per-tool fail-open-or-closed decisions, all in one place.
Guest: Exactly, and none of those items is new information at this point, which is the point — it’s just making sure nothing we discussed stays theoretical. There’s a running lab implementation with capability scoping, schema validation, token buckets, approval gates, and the audit log wired together, deliberately protocol-independent, so the same checklist applies whether the calls come over MCP or something else entirely.
9. Seeing it run: the reference implementation and the interview test
Host: So if someone wants to see this rather than just hear about it, there’s the policy-gated-tool-runtime lab — walk me through what actually running it proves that the discussion alone doesn’t.
Guest: It shows scope and schema catching genuinely different failures, rate limits keyed per tool so a search and a fund transfer never share a budget, and approval gates that block with a hard timeout instead of hanging forever. Every denial lands in the audit log as a first-class event, and because the clock is mocked, all forty-four tests covering every stage run in under a second. And critically it’s not an MCP implementation — it’s the policy layer that sits in front of whatever protocol shows up, which is why MCP dropping its handshake and SSE transport this year didn’t touch a line of it.
Host: Which leaves one thing worth being able to defend cold if someone pushes back on this in an interview or a design review: log the denials, not just the successes, because a system that only records what worked can’t show you the caller quietly probing a door it was never given a key to — and that’s exactly the signal you need after something’s already gone wrong. That’s the episode.
Not covered
The planner wanted these and found nothing in the source to support them:
- A live walkthrough of an actual production incident caused by a policy-gating failure
- Comparison with specific competing commercial agent-platform products
- Detailed cryptographic mechanics of mTLS or JWT verification
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Problem
Section titled “Problem”A model deciding to call a tool is a request, not an authorization. Treating the two as the same thing is how an agent platform ends up issuing refunds, deleting records, or emailing customers because a model was persuaded to — by a user, by retrieved content, or by a tool description written by someone else.
The naive architecture wires the model’s tool-call output straight into a dispatch table. That is a system whose effective permission set is whatever the model can be talked into, and whose blast radius is every capability any agent was ever given.
What is needed is an enforcement layer that sits between the decision and the action, and answers five separate questions in a fixed order. They are separate because each fails differently, and collapsing any two of them produces a system that is permissive in one direction and brittle in the other.
Requirements
Section titled “Requirements”- Capability scope. A caller may only invoke tools it was explicitly granted.
- Argument validation. Every call is checked against the tool’s schema server-side, regardless of what the model was shown.
- Per-tool quota. Limits are owned by the tool, not shared across the platform.
- Human approval for high-risk actions, with a bounded wait and a defined timeout behaviour.
- Complete audit. Every attempt is recorded with its outcome, including — especially — refusals.
- Attributable identity. Every decision is tied to a verified caller.
Constraints
Section titled “Constraints”- The model is not a trust boundary. Tool descriptions and schemas shown to a model are hints that shape its output; they constrain nothing. Every check must hold against an adversarial caller as well as a confused one.
- Tool descriptions are untrusted input. They enter the model’s context, so a compromised or hostile tool source is a prompt-injection vector — the same risk class as user input, arriving through metadata.
- Approval blocks a request. Any human-in-the-loop step introduces a wait bounded only by human availability, so it must have a timeout and a defined outcome when it fires.
- Identity must be verified, not asserted. Every control below is scoped by caller, so a forgeable caller identity silently voids all of them at once.
Request Flow
Section titled “Request Flow”flowchart TB
C["Agent tool call<br/>(caller identity, tool, arguments)"] --> S
S{"1. Scope check<br/>is this caller granted<br/>this tool at all?"}
S -->|no| D1["DENIED: out of scope"]
S -->|yes| V
V{"2. Schema validation<br/>is this specific call<br/>well-formed?"}
V -->|no| D2["DENIED: invalid arguments"]
V -->|yes| R
R{"3. Rate limit<br/>per tenant AND per tool,<br/>never a shared budget"}
R -->|exhausted| D3["DENIED: rate limited"]
R -->|allowed| A
A{"4. Approval gate<br/>high-risk tools only,<br/>bounded wait"}
A -->|denied| D4["DENIED: rejected by operator"]
A -->|"timed out"| D5["DENIED: approval timeout"]
A -->|approved / not required| H
H["5. Handler executes"] --> OK["ALLOWED: result returned"]
D1 --> AUD
D2 --> AUD
D3 --> AUD
D4 --> AUD
D5 --> AUD
OK --> AUD
AUD["Append-only audit log<br/>every attempt, allowed or denied"]The order is load-bearing, not incidental:
Scope first, because it is the cheapest check and the most absolute. A caller with no grant on a tool should never reach argument validation — both to avoid the work, and to avoid schema error messages that describe a tool it is not entitled to know exists.
Validation before quota, so a malformed call cannot consume a well-behaved caller’s budget. Inverting these lets an attacker exhaust a victim’s quota with garbage.
Approval last, because it is the only stage that blocks on a human. There is no point paying that latency for a call that would have failed a cheaper check anyway.
Failure Modes
Section titled “Failure Modes”Trusting the schema the model was given
The schema in a tool listing shapes what a well-behaved model emits. It guarantees nothing about what arrives. A server that skips validation because “the model was given the schema” is one hallucinated argument — or one adversarial client — from executing a call it never intended to accept.
Tool poisoning through descriptions
Tool descriptions are part of the model’s context, so a hostile or compromised tool source can write a description crafted to steer behaviour well beyond what the tool does. This is prompt injection arriving through metadata rather than user input, and it bypasses every input filter aimed at the user’s message.
Approval with no timeout
An approval queue without a bounded wait is an availability incident waiting for the first unavailable operator. Requests pile up holding connections and context; the system degrades in a way that looks like a backend outage rather than a policy stall.
Shared quota across tools of different consequence
A cheap read and an irreversible write drawing on one budget means read traffic can exhaust the allowance protecting the dangerous operation — or, worse, that raising the limit to accommodate reads quietly raises it for refunds too.
An audit log that records only successes
Audits exist to answer what was attempted. A log of successful calls cannot show that an agent repeatedly tried a capability it was never granted, which is the single clearest signal of either a compromised caller or a badly scoped grant.
Caller identity from an unverified header
Every control here is scoped by caller. If that caller is a header any client can set, scope, quota, approval routing, and audit attribution are all decorative simultaneously — and the system looks correct in every test that does not try to forge it.
Scaling
Section titled “Scaling”- Quota state is the coordination point. In-process token buckets multiply a caller’s effective
limit by the replica count. A shared atomic limiter (a Redis Lua script, as in
labs/async-ai-gateway) is what makes the limit mean one thing. - Approval is a stateful island in a stateless service. Pending approvals must outlive the replica that accepted them, or a deploy silently drops every in-flight request awaiting a human.
- Audit write volume grows with attempts, not successes, and refusal storms are exactly when the log matters most — so the audit path must not be the thing that fails under attack.
- Registry reads are hot and change rarely, which makes them cacheable, and makes cache invalidation the mechanism by which a revoked grant actually takes effect.
- Schema validation cost scales with payload size, so a large-argument tool deserves a size cap before the validator, not after it.
Security
Section titled “Security”- Verify caller identity from a signed token or mTLS-derived principal, never a bare header.
- Validate every argument server-side, against the server’s own copy of the schema.
- Treat tool descriptions as untrusted, and pin provenance for any tool whose description you did not write.
- Scope grants to least privilege, per caller and per tool, and make revocation take effect without a redeploy.
- Require approval for irreversible or high-cost actions as a property of the tool, not a configuration flag that defaults to off.
- Make the audit log append-only and separately access-controlled — a caller able to edit the record of its own refusals has defeated the point.
- Record the arguments that were refused, subject to redaction, since a refusal without its input is not investigable.
Trade-offs
Section titled “Trade-offs”Fail-closed vs. fail-open when the policy store is unavailable
Failing closed protects the capability surface and turns a policy-store outage into a full agent outage. Failing open preserves function and means an outage silently disables authorization. For tools with irreversible effects, closed is the only defensible default; for read-only tools, open with loud alerting is a legitimate choice. Deciding per tool rather than globally is what makes the trade honest.
Approval timeout: deny or escalate
Denying on timeout is safe and turns operator unavailability into user-visible failure. Escalating keeps the request alive and extends the wait, which is only useful if the escalation target is genuinely more available. The wrong answer is no timeout at all, which converts a policy decision into an unbounded hang.
Per-tool grants vs. role-based grants
Per-tool grants are precise and audit cleanly, and grow unmanageable as tools multiply. Roles compress the grant matrix and introduce the classic problem of a role accumulating capabilities nobody intended. Roles composed of explicit tool grants — rather than roles as opaque labels — keeps auditability while bounding the growth.
Enforcing in the runtime vs. at the protocol boundary
Enforcing inside the tool runtime means the checks apply however the call arrives — MCP, HTTP, or an in-process interface — which is why this architecture is deliberately protocol-independent. Enforcing at a protocol gateway is easier to deploy in front of servers you do not control, and gets bypassed by any path that does not traverse it.
- Approval latency is the dominant cost for gated tools, and it is human-scale. Anything routed through approval should be genuinely consequential, or the gate trains operators to rubber-stamp.
- Audit storage grows with attempt volume, including refusals during an attack — the worst time to be rate-limited by your own logging bill.
- Validation is cheap relative to what it prevents, but a large-payload tool can make it measurable; cap sizes rather than skip checks.
- Per-tool quotas are a spend control, not only an abuse control, when tools call paid APIs.
Observability
Section titled “Observability”- Refusals by reason, per caller, per tool. Scope denial, schema failure, quota exhaustion, approval denial, and approval timeout are five different operational problems that look identical in an aggregate “error rate”.
- Approval queue depth and wait-time distribution, with the timeout marked. If the p95 wait is near the timeout, the gate is about to start denying legitimate work.
- Quota exhaustion rate per tool, which distinguishes a caller hitting a fair limit from a limit set wrong.
- Schema failure rate by tool, a leading indicator that a tool’s schema and its handler have drifted apart.
- Audit write failures alerted on, because an enforcement layer that stops recording is worse than one that stops enforcing — it fails silently.
Production Deployment
Section titled “Production Deployment”Before real traffic
- Caller identity is verified cryptographically, not read from an unverified header.
- Arguments are validated server-side against the server’s own schema copy.
- Rate limits are keyed on
(caller, tool)and enforced through shared atomic state, not per-replica buckets. - Every approval-gated tool has a timeout, and the timeout’s outcome is a deliberate choice recorded per tool.
- Pending approvals survive a replica restart.
- The audit log is append-only, separately access-controlled, and alerted on write failure.
- Refusals are recorded with reason and redacted arguments.
- Grant revocation takes effect without a redeploy, and that path is tested.
- Fail-open versus fail-closed is decided per tool, with irreversible tools closed.
Hands-on Lab
A running implementation of this pipeline: capability scoping, JSON Schema validation, per-tool token buckets, approval gates with a bounded wait, and an append-only audit log. Deliberately protocol-independent — the same checks apply whether calls arrive over MCP or anything else. Read the lab documentation →
labs/policy-gated-tool-runtimeproduction-shaped
Interview Questions
Section titled “Interview Questions”In what order do you run scope, validation, quota, and approval — and why that order?
Scope first: cheapest, most absolute, and it prevents schema errors disclosing tools the caller should not know exist. Validation next, before quota, so malformed calls cannot spend a well-behaved caller’s budget. Approval last, because it is the only stage that blocks on a human and there is no point paying that latency for a call a cheaper check would have rejected. Being able to defend an ordering — rather than listing the controls as an unordered set — is the point of the question.
The model was given the tool's JSON Schema. Why validate arguments again server-side?
Because the schema shapes what a well-behaved model emits and constrains nothing about what arrives. The model can hallucinate a malformed argument, and a client that is not the model at all can send anything. The schema is a hint to the caller; validation is the server’s own guarantee.
Where does prompt injection enter a tool platform, other than user input?
Tool descriptions. They are part of the model’s context, so a hostile or compromised tool source can write one crafted to steer behaviour beyond what the tool does — bypassing every filter aimed at the user’s message. It is why tool provenance matters and why a platform should not treat third-party tool metadata as trusted.
What breaks if the approval queue has no timeout?
Availability. The wait becomes bounded only by human presence, so requests accumulate holding connections and context, and the first unavailable operator turns a policy control into an outage that presents as a backend failure. The timeout is required; whether it denies or escalates is the actual design decision.
Why must an audit log record denials, not just successful calls?
Because the question an audit answers is what was attempted. A success-only log cannot show an agent repeatedly probing a capability it was never granted — the clearest available signal of a compromised caller or a mis-scoped grant. It is also the evidence you need after an incident, when the interesting events are precisely the ones that did not happen.
Your policy store is down. Do you fail open or closed?
Per tool, not globally. Irreversible or high-cost tools fail closed — an authorization outage must not become an authorization bypass. Read-only tools can fail open with loud alerting, trading a bounded exposure for continued function. A single global answer is the sign of a policy that was never really designed.