Skip to content

Policy-Gated Tool Runtime

Listen to this page5:26
Read the transcript

1. Five separate questions, not one filter

Host: There’s a moment in a lot of agent systems where the model decides to call a tool, and everyone quietly treats that decision as the authorization. Today we’re pulling apart a runtime that refuses to make that leap — it says deciding to call and being allowed to call are two entirely different events, gated separately. Walk me through what actually sits between the request and the action.

Guest: There are five checks, and the whole point is that none of them can cover for another. Scope answers whether this caller may use this tool at all. Schema validation answers whether this particular call is well-formed. Rate limiting answers whether this tool, for this tenant, has budget left right now. Approval answers whether a human needs to sign off before anything happens. And audit logging answers what actually happened, including every time the answer was no. Collapse scope and schema into one check, for instance, and you get something too permissive in one direction and too brittle in the other.

Host: Give me a concrete case where keeping them separate actually changes behavior — not just conceptually cleaner, but a design decision that would break if you merged them.

Guest: Rate limits are the clearest one — they’re keyed on tenant and tool together, not on the gateway as a whole, because a read-only search and a fund-moving action sharing one budget is a real hazard. And approval gates block with a timeout instead of an unbounded queue, since a queue with no timeout is just an outage waiting for one unavailable operator. Denials get logged as first-class events too, so a tool nobody ever approves shows up in the audit trail instead of vanishing into application logs — and all of that is tested deterministically, with a mocked clock making token-bucket refills and timeouts instant, so forty-four tests across every enforcement stage run in under a second.

2. Why the checks run in this exact order

Host: So the five checks aren’t just five boxes to tick — the sequence itself is doing work. Why does scope have to be first, before anything else even looks at the call?

Guest: Because it’s the cheapest possible rejection and the most absolute one. If a caller has no grant on a tool, you want them bounced before they ever reach argument validation — otherwise you’re doing schema work for someone who shouldn’t even learn the tool exists, and a validation error message can leak that existence. Scope first means an unauthorized caller gets a flat no, not a clue.

Host: And that same logic carries through the rest of the chain — validation before rate limiting, approval dead last?

Guest: Exactly, and it’s worth stating plainly: validate-then-authorize and authorize-then-validate are not the same system wearing different clothes, they disclose different information and spend different people’s budget. Rate limiting sits after validation so a malformed, garbage call can’t burn down a well-behaved caller’s quota. And approval is last because it’s the only stage that pays human latency — there’s no reason to wake up an approver for a call that a cheaper check would’ve killed anyway.

3. Protocol-independent by design — and why that held up under MCP’s rewrite

Host: So where does this actually live relative to something like MCP? Because we’ve been describing five checks and an order, but I haven’t heard the word protocol once.

Guest: That’s deliberate — this lab isn’t an MCP implementation, it’s the policy layer any MCP server would sit behind, and it doesn’t care whether the call arrived over MCP, a plain HTTP API, or an in-process tool interface. The proof is the 2026-07-28 revision, which tore out MCP’s initialize handshake and protocol-level sessions and deprecated the old HTTP+SSE transport entirely, invalidating a huge amount of 2025-era MCP code. None of that touched this pipeline, because validation, authorization, rate limiting and approval were never protocol concerns — wiring this up to a real MCP host just means putting an SDK in front of the gateway’s call method, with zero change underneath.

4. Where the runtime still leans on trust — identity

Host: So where’s the honest crack in all this? Every check you’ve described is scoped by caller, so how solid is that caller identity actually?

Guest: Not solid at all, honestly — it’s a declared header, x-agent-id, not a cryptographically verified credential, so the gateway trusts whoever sets that header rather than proving who’s behind it. That’s fine while each call maps cleanly to one human request, but the moment an agent is chaining actions across a session on someone else’s behalf, that one-to-one mapping breaks and the header stops meaning what you want it to mean — which is exactly the point where you’d need real per-request identity, not a static label, and exactly when an agent platform becomes worth building rather than a lab. Don’t take my word for any of this, though — spin it up, hit search_docs and issue_refund with that header, then pull the audit log and watch what got allowed, what got blocked pending approval, and what got refused outright, because the log is where the pipeline actually tells you the truth about itself.

Not covered

The planner wanted these and found nothing in the source to support them:

  • A live walkthrough of the async-ai-gateway’s Redis-backed limiter replacing this lab’s in-process buckets
  • A deep dive into MCP’s Multi-Round-Trip Requests or requestState signing, since this lab deliberately doesn’t implement MCP
  • Comparing this runtime’s approval gate to the agent loop’s step/cost budget mechanics in Module 5

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

A model deciding to call a tool is a request, not an authorization. This lab implements the layer that turns one into the other: capability scoping, argument validation, per-tool rate limiting, human approval for high-risk actions, and an audit log that records every attempt — including the ones that were denied.

Source: labs/policy-gated-tool-runtime

  • Scope and schema as separate questions. A capability scope answers “may this caller use this tool at all”; JSON Schema validation answers “is this particular call well-formed”. Neither substitutes for the other, and conflating them produces a system that is permissive in one direction and brittle in the other.
  • Rate limits owned by the tool, not the gateway. A read-only search and a fund-moving action must not share a budget. The limiter is keyed on (tenant, tool).
  • Approval gates with a bounded wait. High-risk tools block until an operator decides — or until the timeout fires. An approval queue without a timeout is an outage waiting for the first unavailable operator.
  • Denials as first-class audit events. Every attempt is recorded with its outcome, so a tool nobody ever approves is visible in the audit log rather than buried in application logs.
  • Deterministic tests for time-dependent policy. A mocked clock makes token-bucket refill and approval timeout instant, so 44 tests covering every enforcement stage run in under a second.

The pipeline is not five interchangeable filters. The order is chosen so that each stage is cheap relative to the one after it, and so that no stage leaks information a caller has not earned.

Scope is checked first because it is the cheapest and most absolute: a caller with no grant on a tool should never reach argument validation, both to avoid the work and to avoid schema error messages that describe a tool they are not entitled to know exists. Rate limiting sits after validation so that malformed calls cannot consume a well-behaved caller’s budget. The approval gate is last before execution because it is the only stage that blocks on a human, and there is no point paying that latency for a call that would have failed a cheaper check anyway.

This lab is not an MCP implementation. It implements the policy layer an MCP server would sit in front of, and it is deliberately protocol-independent: the same five checks apply whether calls arrive over MCP, a bespoke HTTP API, or an agent framework’s in-process tool interface.

That separation is worth more than it looks. MCP’s 2026-07-28 revision removed the initialize handshake and protocol-level sessions and deprecated the HTTP+SSE transport — a protocol change significant enough to invalidate most 2025-era “MCP-style” code. None of it touches this pipeline, because none of these controls were ever protocol concerns. Wiring the runtime to real MCP hosts means putting an official SDK in front of ToolGateway.call(), with no change to enforcement.

See Module 6: MCP for the current protocol.

Terminal window
cd labs/policy-gated-tool-runtime
python3.12 -m venv .venv
source .venv/bin/activate
pip install -e '.[dev]'
uvicorn tool_gateway.app:app --reload
Terminal window
# A tool inside the caller's granted scope
curl -s -X POST localhost:8000/v1/tools/search_docs/call \
-H 'content-type: application/json' -H 'x-agent-id: research-agent' \
-d '{"arguments": {"query": "retention policy"}}'
# A high-risk tool: blocks pending an operator decision
curl -s -X POST localhost:8000/v1/tools/issue_refund/call \
-H 'content-type: application/json' -H 'x-agent-id: support-agent' \
-d '{"arguments": {"order_id": "42", "amount_cents": 500}}'
# Every attempt, allowed or denied, is in the audit log
curl -s localhost:8000/v1/audit
Terminal window
pytest # 44 tests
ruff check .
mypy src
  1. Scope checks and schema validation answer different questions. Both are necessary; neither substitutes for the other.
  2. Rate limits belong to the tool, not the gateway as a whole — a cheap read and an irreversible action should never share a budget.
  3. Human-in-the-loop approval only works with a bounded timeout and a defined failure mode. Unbounded approval is an availability risk disguised as a safety control.
  4. An audit log that records only successes cannot answer the question audits are actually for: what was attempted and refused.
  5. Binding a call to the caller’s declared identity rather than the end user behind it becomes a real problem once an agent’s actions no longer map one-to-one onto a single human request — which is exactly when an agent platform becomes worth building.