Multi-Tenant MCP Server
Read the transcript
1. The intuitive mistake: a credential in `_meta`
Host: So you’re building a multi-tenant MCP server, and the whole spec under 2026-07-28 tells you every request is self-describing — protocol version, client info, capabilities, all riding along in this metadata field on every single call. So the natural move is, well, if identity travels per-request anyway, just drop the tenant token in there too. It feels like the design is basically inviting you to do that.
Guest: It really does, and that’s exactly why this lab starts with that mistake instead of skipping to the fix. You put the tenant credential in that metadata field alongside the other protocol fields, and you write your authorization check against it. Then you call a single tool once, from application code, and the server rejects its own client.
Host: Wait, rejects its own client from one call? Walk me through how that even happens.
Guest: The SDK isn’t honest with you about how many requests it’s actually making. Calling a tool triggers an internal follow-up request to check the result against the tool’s schema, and that follow-up carries the SDK’s own protocol stamp but has no slot for your application credential — there’s literally no parameter to pass it through. So the server sees an internal call with no tenant token and rejects it, and the tempting patch is to just exempt that one call type from auth — which quietly reopens the exact hole you built the credential check to close.
2. The fix: authenticate the transport, separate filtering from authorization
Host: So if the credential can’t live in the payload the SDK controls, where does it actually go? What’s the fix that survives the SDK making calls behind your back?
Guest: You move it down to the transport itself — authentication that reads the Authorization header on every single request the connection carries, before any handler even runs. That internal schema-check request still has no idea it needs a tenant token, but it doesn’t need one, because the header is already there on the wire regardless of who inside the SDK issued the call. Unauthenticated gets a 401, invalid gets a 401, and that’s asserted against the real HTTP app, not mocked away.
Host: Okay, so the credential problem is solved. But doesn’t filtering tools/list per tenant already give you the isolation you need?
Guest: That’s the second trap — filtering the list is discovery, not security, because a caller can always name a tool it never saw listed, so tools/call has to be authorized independently or you’ve just published a UI convention. And the refusal for that has to come back byte-identical to an unknown-tool error, because the moment forbidden and nonexistent look different, a tenant can probe every other tenant’s capabilities one call at a time.
3. The discovery boundary and a leak class that didn’t used to exist
Host: Okay, but then server/discover itself has no credential at all — isn’t that the exact hole you just spent five minutes closing? Why is it fine to leave that one endpoint wide open to anyone who can reach it?
Guest: Because requiring a credential there would deadlock bootstrapping — a client would need to authenticate before it could even learn how to authenticate. It’s only safe because this particular server’s discover response was verified to carry nothing but capability flags and supported protocol versions, no tool or resource names, and that’s a property you check per server, not something the protocol guarantees for you. Which is exactly why the newer revision added a one-word setting on tenant-scoped listings — if a tenant’s filtered tools/list response comes back marked with that setting in the wrong state, every cache on the path, client, gateway, CDN, is told it’s fine to hand that list to a different tenant, and the server itself never sees the leak because it behaved correctly the whole time.
4. Proving it, running it, and the general lesson
Host: So walk me through actually proving this instead of just asserting it. What does running the lab look like, and where does it actually put these guarantees to the test?
Guest: You stand up the server, curl it with no Authorization header, and you get a 401 straight from the transport before any application code runs — that’s the deadlock case turned into a passing test instead of a hoped-for property. Send it a real bearer token and you get a real tools/list; swap tenant tokens and the count changes, two tools for one tenant, three for the other, because refunds are scoped. Then eleven tests run over the actual Streamable HTTP transport, not a mock, including one that round-robins a multi-step interaction across replicas to prove there’s no session affinity holding it together — statelessness isn’t a design intention, it’s something a test either passes or fails. And that’s the part that generalizes way past MCP: any SDK that reaches out on its own — refreshing a token, revalidating a connection, retrying a call you didn’t initiate — will silently miss whatever credential scope you invented at your own call sites, which is exactly why this stops being a server-specific fix and becomes a line item on a principal-level checklist.
Host: Which is the real takeaway, I think — the mistake was never really about MCP or the meta field, it’s the far more general trap of an SDK doing work on your behalf that your security model forgot to account for. Check where your credential actually has to live, check what your framework does when you’re not looking, and write the test that proves it rather than the one that assumes it. That’s where we’ll leave it — thanks for walking through all of this.
Not covered
The planner wanted these and found nothing in the source to support them:
- A deep comparison of this server’s enforcement model against the policy-gated-tool-runtime’s five-stage pipeline
- Cost and observability metrics specific to running this server at scale (only briefly touched via the architecture companion, not the lab itself)
- A full walkthrough of migrating from Dynamic Client Registration to Client ID Metadata Documents
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Serving many tenants from one MCP server means one tenant must never discover, call, or infer
another’s capabilities. This lab implements that on the official SDK, speaking the 2026-07-28
protocol over Streamable HTTP, with every test driven through the SDK’s own client rather than an
in-process shortcut.
Source: labs/multi-tenant-mcp-server
Request path
Section titled “Request path”flowchart TB
C["MCP host<br/>Authorization: Bearer tok-globex"] --> LB["Load balancer<br/>no session affinity required"]
LB --> R1["Replica A"]
LB --> R2["Replica B"]
LB --> R3["Replica C"]
R1 --> TV{"TokenVerifier<br/>runs before any handler"}
TV -->|"invalid or absent"| E["401 invalid_token"]
TV -->|"valid"| MW
subgraph MW["Tenancy middleware"]
direction TB
S{"tenant-scoped method?"}
S -->|"no — server/discover"| PASS["serve unauthenticated:<br/>capability flags only,<br/>no tenant data"]
S -->|yes| A{"method"}
A -->|"tools/list"| F["filter to granted tools<br/>cacheScope: private"]
A -->|"tools/call"| G{"granted?"}
G -->|no| D["refusal identical to<br/>a genuinely unknown tool"]
G -->|yes| H["run handler as this tenant"]
end
MW --> INT["SDK-internal calls<br/>(validate_tool_result → tools/list)<br/>carry the same header"]The mistake this lab is built around
Section titled “The mistake this lab is built around”This was built the wrong way first, and the wrong way is the intuitive one.
Under 2026-07-28 every request is self-describing: protocol version, client info, and capabilities
all travel in _meta on every call. So putting the tenant credential in _meta alongside them
looks natural — per-request identity for a per-request protocol.
It does not work, and the failure is instructive. An SDK makes protocol calls that application
code never issues. Concretely, calling a tool the client has not listed yet makes call_tool()
run validate_tool_result(), which issues its own tools/list to fetch the output schema. Both
requests carry the SDK’s own protocol _meta stamp — it adds that to everything it sends — but only
the first carries the application’s: the second is sent by the SDK itself, so a tenant credential
cannot reach it, even though list_tools() accepts meta= when the application is the caller. A server authorizing on _meta therefore
rejects its own client’s internal call. Worse, the obvious workaround — exempt tools/list from
authentication — reopens precisely the hole the design was closing.
The fix is to authenticate on the transport, via the SDK’s TokenVerifier seam, so the
Authorization header covers every request the transport carries regardless of who issued it. That
property has its own test:
async def test_the_credential_covers_the_sdks_own_internal_calls(server: MCPServer) -> None: recorder = WireRecorder() app_meta = cast(RequestParamsMeta, {"app.example/tenant": "acme"}) async with connect(server, ACME_TOKEN, recorder=recorder) as acme: result = await acme.call_tool("search_docs", {"query": "retention"}, meta=app_meta)
assert not result.is_error call = next(r for r in recorder.requests if r.method == "tools/call") hidden = next(r for r in recorder.requests if r.method == "tools/list") assert "app.example/tenant" in call.meta_keys assert "app.example/tenant" not in hidden.meta_keys assert "io.modelcontextprotocol/protocolVersion" in hidden.meta_keys assert hidden.authorization == f"Bearer {ACME_TOKEN}"It asserts on what the server actually received. An earlier version asserted only that the call succeeded — which it would equally have done had the internal request never existed.
What it demonstrates
Section titled “What it demonstrates”- Transport-level authentication, checked before any handler runs — 401 without a credential, 401 with an invalid one, both asserted against the real HTTP app.
- Tokens bound to this server. Each token records the resource it was issued for (RFC 8707), and a genuine tenant token minted for a different server gets a 401.
- Per-tenant discovery. Each tenant’s
tools/listis filtered to its own grants. - Authorization separate from filtering.
tools/callis checked independently, because a caller can always name a tool it never listed. A server that only filters the listing has published a UI convention, not a security boundary. - Refusals that leak nothing. An ungranted call returns a result byte-identical to a genuinely unknown tool. A distinguishable refusal is an enumeration oracle — a tenant that can tell “forbidden” from “nonexistent” can map every other tenant’s capabilities one call at a time.
cacheScope: privateon tenant-scoped listings.publicwould tell every cache on the path — client, gateway, CDN — that one tenant’s filtered list may be served to another. This is a new way to leak tenant data; the previous protocol revision had nocacheScopeto get wrong.- Statelessness proven, not assumed. A test alternates tenants across three connections to one server object and asserts no identity carries over.
What re-verification found
Section titled “What re-verification found”On 2026-10-08, re-verifying this page against mcp 2.3.0 turned up three faults in the lab itself.
None of them was visible from a green CI run, which is the reason to record them.
- The tests ran the wrong protocol. The harness connected through the low-level
ClientSession, whoseinitialize()performs the legacy handshake and negotiates2025-11-25. Every test passed, and none exercised2026-07-28. The harness now uses the SDK’sClient, and a test asserts on the wire that the first request isserver/discoverand noinitializeis sent. - The central test inferred instead of observing. It is the one above, now asserting on what the server received.
- It accepted tokens minted for other servers. Nothing checked a token’s resource: the
verifier did not, and the SDK did not until 2.2.0, and then only when asked —
validate_token_resourcedefaults off in 2.x and on in 3.0. The lab now sets it, and the SDK floor is 2.2.0.
Each new test was checked against the fault it exists to catch: pointing the harness back at
ClientSession, or turning the resource check off, fails exactly the test written for it.
Where the authentication boundary sits
Section titled “Where the authentication boundary sits”server/discover is served without a credential, and that is deliberate rather than an
oversight. Requiring one would deadlock bootstrapping: a client would need to authenticate before it
could discover how to authenticate.
It is safe only because this server’s discover response was verified to carry capability flags and
supported protocol versions — no tool or resource names. That is a property of this server, not a
general guarantee. A server that put tenant-scoped data in discover would have to move it behind the
boundary, which is why the method list is explicit in middleware.py rather than inferred.
Run it
Section titled “Run it”cd labs/multi-tenant-mcp-serverpython3.12 -m venv .venvsource .venv/bin/activatepip install -e '.[dev]'python -m mcp_tenancy.app# No credential: rejected by the transportcurl -s -o /dev/null -w '%{http_code}\n' -X POST localhost:8000/mcp \ -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \ -H 'Mcp-Method: tools/list' \ -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}' # 401
# With a tenant credentialcurl -s -X POST localhost:8000/mcp \ -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \ -H 'Mcp-Method: tools/list' -H 'Authorization: Bearer tok-globex' \ -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}' # 200, two toolsSwap tok-globex for tok-acme and the same call returns three tools — issue_refund is granted
to acme only.
Verify it
Section titled “Verify it”pytest # 11 tests, all over the real Streamable HTTP transportruff check .mypy srcPrincipal-level discussion points
Section titled “Principal-level discussion points”- A credential has to live where it covers calls your own code does not make. SDKs issue revalidations, refreshes, and retries on your behalf; anything scoped to application-issued requests will have gaps you cannot see from your own call sites.
- Filtering a listing is discovery, not authorization. Both are needed, and conflating them produces a server that looks isolated in a UI and is not.
- A refusal that differs from “not found” is an enumeration oracle. Error text is part of the security surface.
cacheScopemade caching a tenancy concern. A protocol revision can introduce a leak class that did not previously exist, which is the strongest argument for tracking protocol freshness deliberately rather than by editorial habit.- Statelessness moved authentication from once-per-session to once-per-request. That is more work per call and strictly better operationally — no affinity, no session store, no drain-on-deploy.
Related
Section titled “Related”- Architecture: Enterprise MCP Platform — the design-review companion: constraints, scaling, cost, and a pre-traffic checklist.
- Module 6: MCP — the protocol, the
2026-07-28changes, and the credential-placement lesson this lab produced. labs/policy-gated-tool-runtime— the enforcement pipeline that would sit behind this server’s tool calls.labs/async-ai-gateway— theproduction-readyreference this lab is measured against.