Architecture: Enterprise MCP Platform
Read the transcript
1. Why one server, many tenants, and why now
Host: So let’s start with the problem that actually forces this whole redesign conversation. Once more than one team inside a company wants an MCP server, the naive answer is one server per team, per integration — and that sounds fine until you’re the one operating it. Suddenly every team’s server is its own deployment, its own credential, its own patch cycle, and it just doesn’t scale.
Guest: Right, so the obvious fix is consolidation — one server, many tenants — but that doesn’t make the isolation problem go away, it just relocates it. Instead of isolation being an infrastructure property, where each team’s server is physically separate and a mistake is contained by the deployment boundary, isolation becomes a property of your code. One missing check in that shared server and tenant A’s tools are visible to tenant B.
Host: And this is exactly the moment the protocol itself changed underneath everyone — the 2026-07-28 revision. Walk me through why that timing matters: it’s not just that consolidation is architecturally harder now, it’s that the protocol you’d consolidate on top of got rewritten in ways that touch identity and caching directly.
Guest: Exactly right — two changes matter here. MCP removed the initialize handshake and the session entirely, so identity can’t be established once per connection anymore; every single request has to prove who it is on its own. And the new cacheScope field means a listing response is now something infrastructure is allowed to store and replay, which opens up a leak class — cached data crossing tenant boundaries — that simply didn’t exist before.
2. The seven requirements a multi-tenant server must satisfy
Host: Okay, so given those two protocol shifts, what actually has to be true for a server to call itself safely multi-tenant? Is there a checklist you hold every implementation against?
Guest: There is, and it’s seven requirements. Per-request identity is the foundation — every request authenticates on its own. Then discovery isolation, so a tenant’s tools/list and resources/list only ever return what that tenant was granted, and invocation authorization, which is separate — a tenant can’t call something even if it never showed up in their listing, and can’t call it even if it somehow did. Then non-disclosure on refusal, meaning a denial can’t leak that the capability even exists, cache safety so nothing is ever marked shareable across tenants, horizontal scale so any replica can serve any request with no sticky affinity, and auditability — every call and every refusal traceable to a specific tenant.
Host: That’s a lot to hold simultaneously — discovery, invocation, and refusal all have to independently agree on the same boundary. Which of these is the one teams most often get wrong first?
Guest: Non-disclosure is a subtle one, because teams nail authorization but then their error message gives away that a restricted tool exists, which is its own leak. The rest of this episode is really just walking through each of these seven one at a time and showing where the naive implementation breaks.
3. The constraints nobody designs for on paper: credentials, discover, wire-level filtering
Host: So before we walk through the seven requirements one by one, you said there are constraints nobody designs for on paper. What’s the first one that trips teams up?
Guest: Credential placement — and it’s counterintuitive because the protocol makes every request self-describing, so putting the tenant token in per-request metadata feels right. But the SDK issues calls your application code never writes, like an internal tools/list to validate an output schema after call_tool, and that internal call has no way to carry your application’s metadata. So a server authorizing on request metadata ends up rejecting its own client’s internal call, and the only fix is putting the credential on the transport’s Authorization header so it covers everything the transport carries, not just what your code explicitly sends.
Host: That’s a rough one to discover in production. What about the other constraints — discovery, filtering, and registration?
Guest: Server discover can’t require authentication at all, because a client needs it just to learn how to talk to you, so whatever’s in that response is public by definition — no tenant secrets there. Filtering is trickier than people expect too, because middleware operates on outbound dicts after serialization, not on your typed result models, so isolation logic has to manipulate wire representations directly. And separately, dynamic client registration is deprecated now — new integrations use Client ID Metadata Documents with issuer validation per RFC 9207, so that’s one more place old assumptions quietly stop holding.
4. Tracing a request through the fleet, and where it goes wrong
Host: So walk me through what actually happens when a request lands. It arrives at some replica with a bearer credential, the token verifier resolves that to a tenant before anything else runs — and then what, every tenant-scoped call gets filtered and authorized separately?
Guest: Right, no credential or a bad one gets a 401 at the transport, before your application logic even sees it. Past that, listings get filtered down to what the tenant is actually granted, and invocations are checked against those same grants independently — filtering and authorization are two separate gates, not one gate wearing two hats. The one deliberate exception is server discover, which bypasses all of it.
Host: That sounds clean on paper. Where does it actually break in practice?
Guest: The expensive one is putting the tenant credential in per-request metadata instead of the transport layer, because the SDK’s own tool-calling implementation fires an internal list-tools call to validate output schema, and that internal call carries no application metadata — so the server locks out its own client, and the tempting fix, exempting that internal listing call from auth, reopens the exact hole you built the isolation to close. Then there’s a one-word cache setting introduced in a later revision, which tells every cache on the path — client, gateway, CDN — that one tenant’s filtered list is fine to serve to another, and the server behaves perfectly correctly while the leak happens entirely in infrastructure it never sees. Add to that filtering the tool listing without authorizing the tool invocation call, which is a UI convention masquerading as a security boundary and demos beautifully right up until someone names a tool directly; distinguishable ‘forbidden’ versus ‘not found’ responses that let a tenant enumerate everyone else’s capabilities one call at a time; and code that caches identity or an in-progress operation per connection, which works fine on one replica in dev and fails intermittently the moment a load balancer is in the picture.
5. The security checklist that closes those holes
Host: So we’ve got the wound list — five ways this thing bleeds. Give me the dressing for each, starting with tokens and that credential-placement trap you just described.
Guest: Verify every token against an authorization server and validate the issuer per RFC 9207, binding the credential to that issuer — static tokens are a lab convenience, not a deployment. And use Client ID Metadata Documents instead of Dynamic Client Registration, since this revision deprecated it. Then, separately, authorize the invocation call independently of the discovery call — filtering a list answers ‘what can this tenant see,’ authorizing a call answers ‘can this tenant do this,’ and treating those as one check produces a server that looks isolated in a UI and is not.
Host: And the enumeration and caching holes — the ones that leak without any code being wrong?
Guest: Make refusals byte-identical to not-found — status code, shape, error text, all of it — because any difference is an oracle a tenant can walk. Set cacheScope private on every tenant-scoped response and actually assert it in a test, not code review, since that’s the setting that told every cache on the path a filtered list was fair game to share. And treat tool descriptions from tenant-supplied sources as untrusted — they land in the model’s context, so the tool-poisoning risk is now an inside-the-platform problem, not just an across-servers one. Last piece: sign, bound, and expire requestState, because it’s client-held state that resumes execution server-side, and if you don’t bind it to the authenticated principal the way the SDK does, you’ve handed out a resumable token nobody’s watching — which is also why you audit refusals as carefully as successes, since a refusal is the only evidence a cross-tenant attempt ever happened.
6. The trade-offs a platform team is actually making
Host: So let’s zoom out from the checklist and talk about why any of this is hard in the first place. It sounds like every decision in this design has a cheaper, less-safe alternative sitting right next to it — and picking the safe one costs something real. What’s the first trade-off a platform team actually has to make?
Guest: Shared server versus per-tenant deployment. A shared server amortizes ops, patching, capacity — a new tenant becomes a config change, not a deployment. But that’s exactly what turns isolation into a code property: a missing check is now a cross-tenant breach, where separate deployments would’ve just contained it. Per-tenant buys you that containment back, but the operational cost grows linearly with every tenant you add, which is the thing consolidation was supposed to fix.
Host: And that same shape repeats — cheap-but-risky versus expensive-but-safe — in the other three axes too?
Guest: Exactly the same shape. Per-request verification costs a check on every call but buys a fleet with no affinity, no session store, no partially-authenticated state — and it’s rarely a close call since token verification is cheap; per-session just reintroduces the stickiness the protocol revision removed. Server-side filtering is the only actually secure option since the host isn’t a trust boundary, but returning everything is simpler and lets one cached response serve every tenant — which is exactly why a public cache scope setting is a shortcut, not a feature. And tenant-level grants are simple to audit while role-level grants match how organizations really delegate, except now that dimension is baked into every check, every cache key, every audit query — and retrofitting it later is expensive precisely because the old shape is already encoded everywhere.
7. Scaling a stateless fleet, and making the edge do work
Host: So if per-request identity kills the session, what does that actually buy the platform team running the fleet, day to day?
Guest: It makes the fleet boring, in the good sense. No session store to shard or fail over, no client affinity to maintain, no draining sessions before a deploy — any instance can serve any request, so you scale it like any other stateless service. On top of that, Mcp-Method is visible without parsing the body, so a gateway can rate-limit tools/call separately from tools/list, or route the expensive operations to dedicated capacity. But caching gets subtle: cacheScope private stops shared caches from cross-serving, but it says nothing about a single client acting for several tenants, so your cache keys have to include the tenant explicitly or you’ll serve tenant A’s list to tenant B. And since hosts list far more than they call, getting ttlMs right on discovery is most of the protocol revision’s performance win — get it wrong and you’re re-fetching on every turn.
Host: And that discovery traffic isn’t just a server load problem — you said tool count is a context cost too. What does that mean in practice for how a tenant’s grant gets sized?
Guest: Right — every tool you list typically lands in the model’s context window, so a tenant with a large grant isn’t just hitting your server harder, they’re paying for it in degraded response quality and token spend before the server even breaks a sweat. That flips how you think about role-level grants: it’s not only an authorization question, it’s a UX and cost budget per tenant. So teams end up tuning grants down to what a role actually needs, partly for security, partly because a leaner tool list makes the model perform better.
8. What it costs and what you have to watch
Host: Okay, so grant size is a cost line, not just a risk line. Walk me through the rest of the bill — where does the money actually go once this thing is running at scale?
Guest: Token verification runs per request, not per connection, so it scales with call volume — cache the verification result by token with a short TTL bounded by the token’s own expiry, or you’re paying that cost on every single call. Getting the discovery cache duration right is a direct saving too, since discovery dominates the request mix and every cache hit is a request the fleet never has to serve. And the one nobody budgets for upfront is audit retention — logging every call and refusal per tenant, kept long enough to matter in an investigation, usually ends up the largest storage line in the whole platform.
Host: So how do you actually see any of that going wrong before it shows up as an incident?
Guest: Every log line and span has to carry a tenant identifier or none of this is answerable at all. From there you watch refusal rate per tenant and per tool — a rise is either a misconfigured grant or an enumeration attempt, and only the per-tool breakdown tells you which. You also split 401s by cause, since absent credentials usually mean a broken client while invalid ones mean expiry or rotation, you track the cache hit ratio on the tool listing call because a low number means clients are ignoring your cache duration setting and the fleet’s eating traffic it shouldn’t, you break latency out by method name since blending cheap listings with expensive calls describes nothing, and you watch protocol version distribution to know when the old clients have actually stopped connecting.
9. The pre-traffic checklist and a working reference
Host: So if I’m a platform team about to ship this, what’s the actual gate before traffic hits it — not the design conversation, the checklist someone signs off on?
Guest: It’s the nine tests we already walked through — that’s the actual sign-off gate, not a restatement. If any of those isn’t a test, it’s a hope.
Host: That’s a good place to leave it. If someone wants this as running code instead of a checklist, where do they look?
Guest: The multi-tenant MCP server lab is exactly that — built on the official SDK, with transport auth, per-tenant discovery, invocation checked separately from listing, refusals that leak nothing, and statelessness proven rather than assumed, each one asserted by a real test against the real HTTP app. It’s the same seven requirements and the same failure modes we’ve walked through this whole episode, just compiled and runnable. Worth reading the discover boundary section especially — it’s the one place where deliberately skipping auth is the correct call, and the lab shows exactly why.
Not covered
The planner wanted these and found nothing in the source to support them:
- A real-world incident walkthrough of a cross-tenant breach in production
- Specific dollar figures or benchmark numbers for token verification or audit storage costs
- Comparison of this platform’s design against a competing vendor’s multi-tenant MCP offering
- Migration playbook for moving an existing per-team-server fleet onto the shared platform
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Problem
Section titled “Problem”Once more than one team wants MCP servers, running one server per team per integration stops scaling: every server is a separate deployment, a separate credential, a separate thing to patch. The obvious consolidation — one server, many tenants — moves the problem into the server, where a single missing check exposes one customer’s tools to another.
The 2026-07-28 protocol revision changes the shape of that problem in two ways that matter
architecturally. Removing the session means identity can no longer be established once per
connection; it has to be re-established per request. Adding cacheScope means a listing response is
now something infrastructure may legitimately store and replay — which creates a leak class that did
not previously exist.
Warning
This page describes the 2026-07-28 protocol. Material written for 2025-11-25 assumes an
initialize handshake and a session that this revision removed, which changes both the identity
model and the deployment topology below.
Requirements
Section titled “Requirements”- Per-request identity. Every request authenticates on its own; nothing is trusted because an earlier request on the same connection was trusted.
- Discovery isolation. A tenant’s
tools/listandresources/listreturn only what that tenant was granted. - Invocation authorization. A tenant cannot call what it was not granted, whether or not it appeared in a listing.
- Non-disclosure on refusal. A refusal must not reveal that the refused capability exists.
- Cache safety. No response is ever advertised as shareable across tenants.
- Horizontal scale without affinity. Any replica serves any request.
- Auditability. Every call and every refusal is attributable to a tenant.
Constraints
Section titled “Constraints”- Credentials must live where the SDK’s own calls carry them. An SDK issues protocol requests
application code never writes — validation, revalidation, retries. Any credential scoped to
application-issued requests has gaps invisible from your own call sites. In practice this means
the transport’s
Authorizationheader, not per-request_meta. server/discovercannot require authentication. A client needs it to learn how to talk to the server, so gating it deadlocks bootstrapping. Everything in that response is therefore public by construction.- Filtering happens below model serialization. Middleware sees outbound dicts, not typed result models, so isolation logic manipulates the wire representation.
- Dynamic Client Registration is deprecated. New integrations use Client ID Metadata Documents, and issuers are validated per RFC 9207.
Request Flow
Section titled “Request Flow”flowchart TB
C["MCP host<br/>Authorization: Bearer tok-globex"] --> LB["Load balancer<br/>no session affinity required"]
LB --> R1["Replica A"]
LB --> R2["Replica B"]
LB --> R3["Replica C"]
R1 --> TV{"TokenVerifier<br/>runs before any handler"}
TV -->|"invalid or absent"| E["401 invalid_token"]
TV -->|"valid"| MW
subgraph MW["Tenancy middleware"]
direction TB
S{"tenant-scoped method?"}
S -->|"no — server/discover"| PASS["serve unauthenticated:<br/>capability flags only,<br/>no tenant data"]
S -->|yes| A{"method"}
A -->|"tools/list"| F["filter to granted tools<br/>cacheScope: private"]
A -->|"tools/call"| G{"granted?"}
G -->|no| D["refusal identical to<br/>a genuinely unknown tool"]
G -->|yes| H["run handler as this tenant"]
end
MW --> INT["SDK-internal calls<br/>(validate_tool_result → tools/list)<br/>carry the same header"]A request arrives at any replica carrying a bearer credential. The token verifier resolves it to a
tenant before dispatch; an absent or invalid credential is rejected at the transport with 401.
Tenant-scoped methods then pass through isolation: listings are filtered to the tenant’s grants, and
invocations are authorized independently against those same grants.
server/discover bypasses that boundary deliberately, and is safe to do so only because its
response carries capability flags and supported protocol versions — never tenant-scoped names. That
is a property to verify per server, not a general guarantee.
Failure Modes
Section titled “Failure Modes”Credentials that the SDK's own requests do not carry
The most expensive mistake available here, because it looks correct in every application code
path. Putting the tenant credential in per-request _meta fails the moment the SDK makes a call
of its own: calling a tool it has not listed yet, call_tool() internally issues a tools/list
to fetch the output schema, carrying no application metadata. The server refuses its own client. The tempting fix — exempt tools/list
from authentication — reopens exactly the hole the design was closing.
cacheScope: public on a tenant-filtered listing
A one-word setting that instructs every cache on the path — client, gateway, CDN — that one
tenant’s filtered tool list may be served to another. It has no local symptom: the server behaves
correctly, and the leak happens in infrastructure the server never sees. This failure mode did not
exist before 2026-07-28 introduced the field.
Filtering the listing and calling it authorization
A caller can name a tool it never listed. A server that filters tools/list but does not check
tools/call has built a UI convention, not a security boundary — and it demos perfectly, because
every client-driven path respects the filter.
Refusals that confirm existence
If “forbidden” is distinguishable from “not found”, a tenant can enumerate every other tenant’s capabilities one call at a time. Error text and status codes are part of the security surface, and this is usually discovered by a security review rather than a test.
Per-connection state under a stateless protocol
Code ported from the previous revision often caches an identity, a capability set, or an in-progress operation for “the connection”. Nothing guarantees the next request lands on the same replica, so the symptom is intermittent and load-dependent — correct on one replica in development, wrong behind a load balancer.
Tenant-scoped data in the discover response
Because server/discover is served unauthenticated, anything that leaks into it is public. A
server that grows tenant-specific capability advertisement has published it to anyone who can
reach the endpoint.
Scaling
Section titled “Scaling”- Statelessness is what makes the fleet ordinary. No session store, no affinity, no drain-on- deploy for session migration. This is the largest operational consequence of the protocol revision and it accrues to platform teams rather than tool authors.
Mcp-Methodenables policy at the edge. Because the operation is visible without parsing the body, a gateway can rate limittools/calldifferently fromtools/list, route expensive operations to dedicated capacity, or shed load by operation class.- Cache keys must include the tenant.
cacheScope: privateprevents shared caches from cross-serving; it says nothing to a single client acting for several tenants, whose own cache keys on the method. Multi-tenant hosts need per-tenant cache partitions. - Discovery dominates request mix. Hosts list far more often than they call. Correct
ttlMsturns most of that traffic into cache hits, and ignoring the field forfeits the revision’s main performance win. - Tool count per tenant is a context cost, not just a server cost. Every listed tool typically enters the model’s context, so large grants degrade the caller’s quality and spend before they strain the server.
Security
Section titled “Security”- Verify tokens against an authorization server, validating the issuer per RFC 9207 and binding the credential to that issuer. Static tokens are a lab convenience, not a deployment.
- Use Client ID Metadata Documents rather than Dynamic Client Registration, which this revision deprecated.
- Authorize invocation independently of discovery, since the two answer different questions.
- Make refusals indistinguishable from “does not exist”, including status code, shape, and text.
- Set
cacheScope: privateon every tenant-scoped response, and assert it in a test rather than leaving it to review. - Treat tool descriptions from any tenant-supplied source as untrusted, since they enter the model’s context — the tool-poisoning risk from Module 6 applies within a platform as much as across one.
- Sign, bound, and expire
requestState. In a Multi-Round-Trip Request it is client-held state that resumes server-side execution; the official SDK binds it to the authenticated principal, and a platform should not do less. - Audit refusals as carefully as successes. Refusals are the security-interesting half, and the only durable evidence of attempted cross-tenant access.
Trade-offs
Section titled “Trade-offs”Shared multi-tenant server vs. one deployment per tenant
A shared server amortizes operations, patching, and capacity across tenants, and makes a new tenant a configuration change rather than a deployment. The cost is that isolation becomes a property of code rather than of infrastructure — a missing check is a cross-tenant breach, where separate deployments would have contained it. Per-tenant deployment buys isolation with operational cost that grows linearly in tenants.
Per-request authentication vs. per-session authentication
Per-request costs a verification on every call and buys a fleet with no affinity requirement, no session store, and no partially-authenticated state to reason about. Per-session amortizes the verification and reintroduces exactly the stickiness the protocol revision removed. Token verification is cheap enough that this trade is rarely close for a network-exposed server.
Filtering listings server-side vs. returning everything and filtering at the host
Server-side filtering is the only option that is actually secure, since the host is not a trust
boundary. Returning everything is simpler and lets one cached response serve all tenants — which
is precisely why it is wrong, and precisely the shortcut cacheScope: public makes available.
Granting tools per tenant vs. per tenant-and-role
Tenant-level grants are simple to model and audit. Role-level grants within a tenant match how organizations actually delegate, at the cost of a second dimension in every check, every cache key, and every audit query. Adding the dimension later is expensive because cache keys and audit history both encode the old shape.
- Token verification runs per request, so its cost scales with call volume rather than connection count. Cache verification results by token with a short TTL, bounded by the token’s own expiry.
- Correct
ttlMsis a direct saving, because discovery dominates the request mix and each cache hit is a request the fleet does not serve. - Per-tenant grant size drives model spend, since listed tools enter the caller’s context on every turn. Over-granting is a bill, not just a risk.
- Audit retention is the quiet cost. Recording every call and refusal per tenant, retained long enough to be useful in an investigation, is usually the largest storage line in the platform.
Observability
Section titled “Observability”- Every log line and span carries a tenant identifier, or nothing here is answerable.
- Refusal rate per tenant, per tool. A rise is either a misconfigured grant or an enumeration attempt, and distinguishing them requires the per-tool breakdown.
401rate split by cause — absent credential versus invalid one. Absent usually means a broken client; invalid usually means expiry or rotation.- Cache hit ratio on
tools/list. A low ratio means clients are ignoringttlMs, and the fleet is serving traffic it should not see. - Latency by
Mcp-Method. Aggregate latency blends cheap listings with expensive calls into a number that describes neither. - Protocol version distribution across callers, which is what tells you when legacy-era clients have actually stopped connecting.
Production Deployment
Section titled “Production Deployment”Before real traffic
- Credentials are verified at the transport, and a test proves the SDK’s own internal calls are covered.
- Tokens are validated against an authorization server with RFC 9207 issuer validation; CIMD replaces Dynamic Client Registration.
tools/callis authorized independently oftools/listfiltering, and a test calls an unlisted-but-existing tool.- Refusals are byte-identical to “unknown”, asserted by test.
- Every tenant-scoped response sets
cacheScope: private, asserted by test. server/discoveris confirmed to contain no tenant-scoped names.- The fleet runs stateless with no session affinity, verified by round-robining a multi-step interaction across replicas.
- Audit records tenant, method, tool, and outcome for every request including refusals.
- Protocol version support is pinned and monitored, with a review trigger when the SDK adds a version.
Hands-on Lab
A running implementation of this architecture on the official SDK: transport-level tenant authentication, per-tenant discovery, invocation authorized separately, refusals that leak nothing, and statelessness proven by test. Read the lab documentation →
labs/multi-tenant-mcp-serverproduction-shaped
Interview Questions
Section titled “Interview Questions”Where does the tenant credential live in an MCP platform, and why?
On the transport — the Authorization header — not in per-request _meta. An SDK issues protocol
calls application code never writes; the Python SDK’s call_tool() internally issues a
tools/list to fetch an unlisted tool’s output schema, carrying no application metadata. A credential in _meta
therefore fails on the SDK’s own calls, and exempting those calls reopens the hole. The general
form: a credential must live where every request the transport sends carries it, including the
ones you did not write.
Your tools/list is filtered per tenant. What does cacheScope have to be, and what happens if it is wrong?
private. public tells every cache on the path that the response may be shared, so a gateway or
CDN can serve one tenant’s filtered list to another. The server itself behaves correctly
throughout — the leak lives in infrastructure the server never sees, which is why this belongs in
a test rather than a review comment. It is also new: the previous protocol revision had no
cacheScope to get wrong.
Is filtering tools/list per tenant sufficient isolation?
No. A caller can name a tool it never listed, so tools/call must be authorized independently.
Filtering is discovery; authorization is a separate decision. A server that does only the first
demos perfectly, because every client-driven path respects the filter.
Why must a refusal look identical to 'unknown tool'?
Otherwise it is an enumeration oracle: a tenant that can distinguish “exists but forbidden” from “does not exist” can map every other tenant’s capabilities one call at a time, which undoes the listing filter entirely. Error text, shape, and status are part of the security surface.
Why is server/discover served without authentication?
Because requiring a credential to discover how to authenticate is a bootstrapping deadlock. That is only safe while the response contains no tenant-scoped data — capability flags and supported protocol versions only. It is a property to verify for your server rather than assume, and a reason to keep the unauthenticated method list explicit rather than inferred.
What did removing the session change about how you deploy this?
It removed session affinity as a deployment requirement, which removes the session store, sticky routing, and session-draining on deploy. The fleet becomes an ordinary stateless HTTP service. The cost is that identity is verified per request rather than once per connection, and any code that implicitly cached per-connection state has to stop.