Skip to content

Multi-Tenant MCP Server

Listen to this page5:50
Read the transcript

1. The intuitive mistake: a credential in `_meta`

Host: So you’re building a multi-tenant MCP server, and the whole spec under 2026-07-28 tells you every request is self-describing — protocol version, client info, capabilities, all riding along in this metadata field on every single call. So the natural move is, well, if identity travels per-request anyway, just drop the tenant token in there too. It feels like the design is basically inviting you to do that.

Guest: It really does, and that’s exactly why this lab starts with that mistake instead of skipping to the fix. You put the tenant credential in that metadata field alongside the other protocol fields, and you write your authorization check against it. Then you call a single tool once, from application code, and the server rejects its own client.

Host: Wait, rejects its own client from one call? Walk me through how that even happens.

Guest: The SDK isn’t honest with you about how many requests it’s actually making. Calling a tool triggers an internal follow-up request to check the result against the tool’s schema, and that follow-up carries the SDK’s own protocol stamp but has no slot for your application credential — there’s literally no parameter to pass it through. So the server sees an internal call with no tenant token and rejects it, and the tempting patch is to just exempt that one call type from auth — which quietly reopens the exact hole you built the credential check to close.

2. The fix: authenticate the transport, separate filtering from authorization

Host: So if the credential can’t live in the payload the SDK controls, where does it actually go? What’s the fix that survives the SDK making calls behind your back?

Guest: You move it down to the transport itself — authentication that reads the Authorization header on every single request the connection carries, before any handler even runs. That internal schema-check request still has no idea it needs a tenant token, but it doesn’t need one, because the header is already there on the wire regardless of who inside the SDK issued the call. Unauthenticated gets a 401, invalid gets a 401, and that’s asserted against the real HTTP app, not mocked away.

Host: Okay, so the credential problem is solved. But doesn’t filtering tools/list per tenant already give you the isolation you need?

Guest: That’s the second trap — filtering the list is discovery, not security, because a caller can always name a tool it never saw listed, so tools/call has to be authorized independently or you’ve just published a UI convention. And the refusal for that has to come back byte-identical to an unknown-tool error, because the moment forbidden and nonexistent look different, a tenant can probe every other tenant’s capabilities one call at a time.

3. The discovery boundary and a leak class that didn’t used to exist

Host: Okay, but then server/discover itself has no credential at all — isn’t that the exact hole you just spent five minutes closing? Why is it fine to leave that one endpoint wide open to anyone who can reach it?

Guest: Because requiring a credential there would deadlock bootstrapping — a client would need to authenticate before it could even learn how to authenticate. It’s only safe because this particular server’s discover response was verified to carry nothing but capability flags and supported protocol versions, no tool or resource names, and that’s a property you check per server, not something the protocol guarantees for you. Which is exactly why the newer revision added a one-word setting on tenant-scoped listings — if a tenant’s filtered tools/list response comes back marked with that setting in the wrong state, every cache on the path, client, gateway, CDN, is told it’s fine to hand that list to a different tenant, and the server itself never sees the leak because it behaved correctly the whole time.

4. Proving it, running it, and the general lesson

Host: So walk me through actually proving this instead of just asserting it. What does running the lab look like, and where does it actually put these guarantees to the test?

Guest: You stand up the server, curl it with no Authorization header, and you get a 401 straight from the transport before any application code runs — that’s the deadlock case turned into a passing test instead of a hoped-for property. Send it a real bearer token and you get a real tools/list; swap tenant tokens and the count changes, two tools for one tenant, three for the other, because refunds are scoped. Then eleven tests run over the actual Streamable HTTP transport, not a mock, including one that round-robins a multi-step interaction across replicas to prove there’s no session affinity holding it together — statelessness isn’t a design intention, it’s something a test either passes or fails. And that’s the part that generalizes way past MCP: any SDK that reaches out on its own — refreshing a token, revalidating a connection, retrying a call you didn’t initiate — will silently miss whatever credential scope you invented at your own call sites, which is exactly why this stops being a server-specific fix and becomes a line item on a principal-level checklist.

Host: Which is the real takeaway, I think — the mistake was never really about MCP or the meta field, it’s the far more general trap of an SDK doing work on your behalf that your security model forgot to account for. Check where your credential actually has to live, check what your framework does when you’re not looking, and write the test that proves it rather than the one that assumes it. That’s where we’ll leave it — thanks for walking through all of this.

Not covered

The planner wanted these and found nothing in the source to support them:

  • A deep comparison of this server’s enforcement model against the policy-gated-tool-runtime’s five-stage pipeline
  • Cost and observability metrics specific to running this server at scale (only briefly touched via the architecture companion, not the lab itself)
  • A full walkthrough of migrating from Dynamic Client Registration to Client ID Metadata Documents

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

Serving many tenants from one MCP server means one tenant must never discover, call, or infer another’s capabilities. This lab implements that on the official SDK, speaking the 2026-07-28 protocol over Streamable HTTP, with every test driven through the SDK’s own client rather than an in-process shortcut.

Source: labs/multi-tenant-mcp-server

This was built the wrong way first, and the wrong way is the intuitive one.

Under 2026-07-28 every request is self-describing: protocol version, client info, and capabilities all travel in _meta on every call. So putting the tenant credential in _meta alongside them looks natural — per-request identity for a per-request protocol.

It does not work, and the failure is instructive. An SDK makes protocol calls that application code never issues. Concretely, calling a tool the client has not listed yet makes call_tool() run validate_tool_result(), which issues its own tools/list to fetch the output schema. Both requests carry the SDK’s own protocol _meta stamp — it adds that to everything it sends — but only the first carries the application’s: the second is sent by the SDK itself, so a tenant credential cannot reach it, even though list_tools() accepts meta= when the application is the caller. A server authorizing on _meta therefore rejects its own client’s internal call. Worse, the obvious workaround — exempt tools/list from authentication — reopens precisely the hole the design was closing.

The fix is to authenticate on the transport, via the SDK’s TokenVerifier seam, so the Authorization header covers every request the transport carries regardless of who issued it. That property has its own test:

async def test_the_credential_covers_the_sdks_own_internal_calls(server: MCPServer) -> None:
recorder = WireRecorder()
app_meta = cast(RequestParamsMeta, {"app.example/tenant": "acme"})
async with connect(server, ACME_TOKEN, recorder=recorder) as acme:
result = await acme.call_tool("search_docs", {"query": "retention"}, meta=app_meta)
assert not result.is_error
call = next(r for r in recorder.requests if r.method == "tools/call")
hidden = next(r for r in recorder.requests if r.method == "tools/list")
assert "app.example/tenant" in call.meta_keys
assert "app.example/tenant" not in hidden.meta_keys
assert "io.modelcontextprotocol/protocolVersion" in hidden.meta_keys
assert hidden.authorization == f"Bearer {ACME_TOKEN}"

It asserts on what the server actually received. An earlier version asserted only that the call succeeded — which it would equally have done had the internal request never existed.

  • Transport-level authentication, checked before any handler runs — 401 without a credential, 401 with an invalid one, both asserted against the real HTTP app.
  • Tokens bound to this server. Each token records the resource it was issued for (RFC 8707), and a genuine tenant token minted for a different server gets a 401.
  • Per-tenant discovery. Each tenant’s tools/list is filtered to its own grants.
  • Authorization separate from filtering. tools/call is checked independently, because a caller can always name a tool it never listed. A server that only filters the listing has published a UI convention, not a security boundary.
  • Refusals that leak nothing. An ungranted call returns a result byte-identical to a genuinely unknown tool. A distinguishable refusal is an enumeration oracle — a tenant that can tell “forbidden” from “nonexistent” can map every other tenant’s capabilities one call at a time.
  • cacheScope: private on tenant-scoped listings. public would tell every cache on the path — client, gateway, CDN — that one tenant’s filtered list may be served to another. This is a new way to leak tenant data; the previous protocol revision had no cacheScope to get wrong.
  • Statelessness proven, not assumed. A test alternates tenants across three connections to one server object and asserts no identity carries over.

On 2026-10-08, re-verifying this page against mcp 2.3.0 turned up three faults in the lab itself. None of them was visible from a green CI run, which is the reason to record them.

  • The tests ran the wrong protocol. The harness connected through the low-level ClientSession, whose initialize() performs the legacy handshake and negotiates 2025-11-25. Every test passed, and none exercised 2026-07-28. The harness now uses the SDK’s Client, and a test asserts on the wire that the first request is server/discover and no initialize is sent.
  • The central test inferred instead of observing. It is the one above, now asserting on what the server received.
  • It accepted tokens minted for other servers. Nothing checked a token’s resource: the verifier did not, and the SDK did not until 2.2.0, and then only when asked — validate_token_resource defaults off in 2.x and on in 3.0. The lab now sets it, and the SDK floor is 2.2.0.

Each new test was checked against the fault it exists to catch: pointing the harness back at ClientSession, or turning the resource check off, fails exactly the test written for it.

server/discover is served without a credential, and that is deliberate rather than an oversight. Requiring one would deadlock bootstrapping: a client would need to authenticate before it could discover how to authenticate.

It is safe only because this server’s discover response was verified to carry capability flags and supported protocol versions — no tool or resource names. That is a property of this server, not a general guarantee. A server that put tenant-scoped data in discover would have to move it behind the boundary, which is why the method list is explicit in middleware.py rather than inferred.

Terminal window
cd labs/multi-tenant-mcp-server
python3.12 -m venv .venv
source .venv/bin/activate
pip install -e '.[dev]'
python -m mcp_tenancy.app
Terminal window
# No credential: rejected by the transport
curl -s -o /dev/null -w '%{http_code}\n' -X POST localhost:8000/mcp \
-H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \
-H 'Mcp-Method: tools/list' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}' # 401
# With a tenant credential
curl -s -X POST localhost:8000/mcp \
-H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \
-H 'Mcp-Method: tools/list' -H 'Authorization: Bearer tok-globex' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}' # 200, two tools

Swap tok-globex for tok-acme and the same call returns three tools — issue_refund is granted to acme only.

Terminal window
pytest # 11 tests, all over the real Streamable HTTP transport
ruff check .
mypy src
  1. A credential has to live where it covers calls your own code does not make. SDKs issue revalidations, refreshes, and retries on your behalf; anything scoped to application-issued requests will have gaps you cannot see from your own call sites.
  2. Filtering a listing is discovery, not authorization. Both are needed, and conflating them produces a server that looks isolated in a UI and is not.
  3. A refusal that differs from “not found” is an enumeration oracle. Error text is part of the security surface.
  4. cacheScope made caching a tenancy concern. A protocol revision can introduce a leak class that did not previously exist, which is the strongest argument for tracking protocol freshness deliberately rather than by editorial habit.
  5. Statelessness moved authentication from once-per-session to once-per-request. That is more work per call and strictly better operationally — no affinity, no session store, no drain-on-deploy.