Module 6: MCP
Read the transcript
1. Before MCP: the N×M integration problem
Host: So let’s set the scene for why MCP exists at all. Before this protocol, if you wanted a model to talk to, say, your filesystem and also GitHub and also some internal database, you were writing a separate bespoke integration for each pairing, and none of it was reusable by anyone else building a similar tool.
Guest: Right, it’s the classic N times M problem — N applications each needing to talk to M tools, and nobody’s integration code transfers to anyone else’s app. Every team reinvents the same plumbing. MCP’s whole pitch is that you write the integration once, as a server, and it works with any host that speaks the protocol.
Host: And that sounds simple enough on paper, but you told me before we started recording that the really instructive part of MCP right now isn’t the basic client-server split — it’s a specific change in a recent revision. What happened?
Guest: The 2026-07-28 revision ripped out the initialize handshake and the protocol-level session entirely, so every request now has to be self-describing. That’s not cosmetic — a stateful protocol forces sticky routing, which is exactly the state-ownership problem that makes services hard to scale horizontally. Watching MCP make that migration is basically a live case study in stateful-to-stateless design, and that’s what this episode is really about.
2. Three roles, one aggregation pattern
Host: Okay so before we go deeper into the statelessness story, let’s set up the shape of the thing. When you say MCP separates three roles, what actually are they?
Guest: Host, client, server. The host is the application the user actually sees — a chat app, an IDE — and it owns the decision to connect to any given server. The client is a connection manager that lives inside the host, one per connected server, handling protocol mechanics so the host doesn’t have to. And the server is a separate small process exposing tools, resources, or prompts — ideally one server per integration, not one server trying to do everything.
Host: That ‘many small servers behind one host’ shape sounds familiar — didn’t we see almost exactly this pattern somewhere else in the handbook?
Guest: Yeah, it’s the same move as the async AI gateway aggregating multiple LLM providers behind one interface, just applied to tools instead of models. The host connects to a bunch of independently maintained MCP servers and presents their combined capabilities to the model as one coherent context — it’s Module 4’s gateway layer idea again: many small, focused things behind one aggregation point.
Host: So the three-role split isn’t just bureaucratic naming — it’s what makes that aggregation possible in the first place.
Guest: Exactly, the split — host, client, server — is what makes that aggregation possible in the first place.
3. Isolation by design: one client per server
Host: So if the host is juggling multiple servers, why not just have one client that talks to all of them? Why insist on a dedicated client per server?
Guest: Blast-radius isolation, mostly. Each client-server connection is independently scoped, so if one server is compromised or just misbehaving, it only poisons its own client, not the whole host. It also means each connection can use its own transport — one server over stdio, another over HTTP — without forcing a lowest-common-denominator protocol on everyone.
Host: Okay, so that isolation is really about containment and flexibility. But does that mean each client is holding onto some persistent session state for its server?
Guest: That’s the thing people assume, and it’s exactly what the newest revision walks back. As of the July twenty-eighth, twenty twenty-six revision, a client is just a connection manager, not a session holder — isolation doesn’t have to mean statefulness, and that distinction is what we need to unpack next.
4. Tools, resources, prompts: who decides?
Host: Okay, before we go further into statelessness, I want to back up to something you mentioned earlier — tools, resources, and prompts. I’ve been treating those as basically three flavors of the same thing, just different data types. Is that wrong?
Guest: It’s wrong in the way that matters most. The real distinction isn’t what kind of data moves — it’s who decides the invocation happens. A tool is called because the model decided to call it, same function-calling mechanics as always, just standardized across servers. A resource is attached because the host application decided to attach it — a file, a database row, a search result — the model doesn’t reach out and grab it unless the host explicitly allows that. And a prompt only fires because a user picked it, usually as a slash command or a menu item.
Host: So it’s a security boundary dressed up as a taxonomy. Which means the same capability could live in different places depending on who you trust to pull the trigger.
Guest: Exactly, and that’s the trade-off worth sitting with. Take ‘read this file.’ Expose it as a resource, and the host or the user controls when that read happens — predictable, bounded context growth. Expose the identical capability as a tool, and now the model decides for itself when to read it, which gives it real autonomy but also a wider action surface and a context window that grows however the model feels like growing it.
Host: So choosing tool versus resource for the same function isn’t a formatting choice, it’s you deciding how much you trust the model with initiative.
Guest: Right — and that decision doesn’t go away just because the transport underneath it changed. Which is a good bridge, actually, because that discovery pattern — list then invoke, tools/list then tools/call, resources/list then resources/read — is exactly what stays stable while everything about session state gets rebuilt underneath it in the new revision.
5. The stateless rewrite: SEP-2575
Host: Okay, so let’s get concrete about what actually got ripped out. Before, you connect, you do this initialize and initialized handshake, the server learns who you are and what you support, and that sets up a session. What replaces that?
Guest: Nothing replaces the handshake, in the sense that there isn’t one anymore. There’s no initialize, no initialized, no protocol-level session sitting on the connection. Instead, three pieces of information that used to travel once — protocol version, client info, client capabilities — now ride in a _meta block on every single request.
Host: Every request. So if I’m the server, I’m re-learning who’s talking to me on every call instead of once at the start.
Guest: Right, and if you want the server’s capabilities up front the way you used to get them for free from the handshake response, you ask explicitly — there’s a server/discover call for that now. I actually checked this against the official Python SDK rather than just taking the spec’s word for it, and it’s exactly as described: a modern connection issues zero initialize calls, the first thing on the wire is server/discover, and the three _meta keys are literally namespaced as io.modelcontextprotocol/protocolVersion, clientInfo, and clientCapabilities.
Host: And that SDK only speaks the new revision, or does it still know the old handshake?
Guest: It keeps 2026-07-28 as the one modern version, but 2024-11-05 through 2025-11-25 are retained as legacy handshake versions — so old clients aren’t stranded, they just get routed down a different, older code path. But for anything built fresh, the handshake is just gone, and that per-request _meta plus server/discover is the whole replacement.
6. Routing headers and multi-round-trip requests
Host: So the handshake is gone, but requests still have to get routed somewhere before anyone looks at the JSON-RPC body. How does that actually work now?
Guest: That’s SEP-2243. Every request over Streamable HTTP carries an Mcp-Method header, and anything that names a specific target — a tool call, a resource read, a prompt fetch — also carries Mcp-Name. A gateway or load balancer can route and meter purely off those headers without ever parsing the body.
Host: That sounds efficient, but also like an obvious place to lie — put one thing in the header and something else in the body.
Guest: Exactly the gap, and the spec closes it explicitly: servers must reject any request where the header and body disagree. If that check is missing, a caller can route as a cheap, permitted operation while the body actually executes something expensive or restricted — and a hand-rolled server is exactly where that check gets skipped. There’s a related seam in SEP-2322 for calls that need more than one round trip — a tool returns input_required with a requestState string, the client gathers input and echoes it back, and the server resumes from there.
Host: So the protocol stopped holding session state, but that state didn’t disappear — it just moved into a token the client is now responsible for carrying.
Guest: Right, and that’s the trust question worth sitting with. requestState is opaque to the client but trusted by the server on return — if it isn’t signed, expiring, and size-bounded, you’ve built a client-controlled input that resumes server-side execution, which is the same mistake as trusting an unsigned cookie for authorization. The spec leaves the encoding open on purpose, but that openness means the security posture is entirely up to whoever implements the server.
7. Caching and what’s being deprecated
Host: Given that requestState is basically a security decision dressed up as a protocol detail, let’s talk about the other half of discovery — caching. What are ttlMs and cacheScope actually doing there?
Guest: They’re the thing that makes discovery cheap instead of just stateless. When a client calls tools/list or does a resources/read, the response now carries a ttlMs — how long that result stays fresh — and a cacheScope that says whether it’s safe to share across users or has to stay pinned to one. It’s modeled directly on HTTP Cache-Control, and if a client actually honors those fields it stops re-fetching the same tool list on every turn, and a gateway can serve one cached copy to many users when the scope allows it.
Host: And if a client just ignores those fields?
Guest: Then you’ve given up the main performance win of this whole revision — you’re paying full discovery cost every call for no reason. Worth pairing that with the deprecation list, since it’s the same ‘this still works but stop building on it’ theme: HTTP+SSE transport, plus Roots, Sampling, Logging, and Dynamic Client Registration are all deprecated now. Roots, Sampling, Logging, and DCR get a twelve-month minimum window before removal eligibility, landing at the first revision on or after 2027-07-28, but HTTP+SSE is on a faster clock — it’s eligible for removal just three months after its deprecating SEP reaches Final, so that’s the one to migrate off first.
8. A minimal tool server, two validations
Host: Let’s actually look at code, since I think the abstraction has been floating a bit. What does a minimal tool server look like, end to end?
Guest: It’s smaller than people expect. tools/list returns one ToolDefinition — a name, a description, and a JSON Schema for its arguments — plus that ttlMs and cacheScope pair we just discussed. Then tools/call is the interesting part: it’s a stateless handler that takes a name, arguments, and headers, and does two checks before it does anything else.
Host: Two checks. Walk me through why there are two, not one.
Guest: First it checks that arguments actually satisfy the schema — is query a non-empty string — even though the server already published that schema in tools/list. That’s defending against the model: it can hallucinate a malformed call, or a client can forge one, so the schema is a hint to the caller, never a guarantee the server gets to skip validation. Second, completely separately, it checks that the Mcp-Method and Mcp-Name headers agree with the body — that’s defending against infrastructure, for the reasons we already went through with routing headers.
9. IDE and GitHub server: stateless in production
Host: Let’s ground this in something concrete. Walk me through an actual coding task where a host is talking to two servers at once.
Guest: Say you’re in an IDE working on a bug fix. The IDE host is connected to a filesystem server that exposes your project as resources, and a remote GitHub server that exposes tools like create_pull_request and search_issues. You attach a specific file as context — that’s a resource, so you or the host chose it, not the model — the model reads it, understands the bug, then once it has a fix, it decides on its own to call create_pull_request. Resource in, tool call out, two different servers, two different trust boundaries, exactly like we’ve been describing.
Host: So that’s the workflow. Now what does statelessness actually buy the team running that GitHub server?
Guest: It turns it into a completely ordinary service to operate. The operator runs three replicas behind a plain load balancer, no sticky sessions, because there’s no session to pin a client to an instance. They use Mcp-Method to rate limit tools/call harder than tools/list since a pull request costs more than a listing. And they can do rolling deploys that just drop connections without anyone’s in-flight work breaking, because there’s no session state living on a specific instance to lose.
Host: And under the old revision, every one of those would have needed session affinity or a shared session store.
Guest: Right, sticky sessions at the load balancer, or shared session storage across replicas, and deploys would need to drain sessions carefully before rotating instances. That’s the biggest operational payoff of the 2026-07-28 revision — it matters more to the platform team running the GitHub server than to whoever wrote the tool handler, because it’s the difference between a bespoke stateful fleet and a service that scales like any other HTTP backend.
10. Security: tool poisoning, requestState, and the credential-placement trap
Host: Okay, scaling is sorted, so let’s talk about what can go wrong once you’re actually running one of these servers in production. Where does the security story start?
Guest: It starts with a mindset shift: tool and resource descriptions are part of the model’s context, so they’re untrusted input just like anything a user types. A malicious server can write a tool description that manipulates model behavior beyond what the tool actually does — same prompt injection risk you’d worry about anywhere else, just arriving through metadata instead of a chat box. So you only connect to servers whose provenance you actually trust, and you validate every tools/call argument server-side no matter what schema the model was handed.
Host: Right, the schema constrains a well-behaved client, not what the server should assume it received. What about that requestState mechanism you mentioned earlier — the thing that resumes execution?
Guest: That’s the other one people underestimate. requestState is server state handed to the client and trusted on return — if you don’t sign it, bound its size, and expire it, you’ve created a client-controlled input that resumes server-side execution, which is the same mistake as trusting an unsigned cookie for authorization. The official SDK actually ships an AES-GCM codec that binds the state to the authenticated principal, which tells you exactly what threat model the spec authors were worried about.
Host: So that covers metadata and state. But you keep coming back to credential placement as the real trap — walk me through the multi-tenant lab mistake.
Guest: Because every request is self-describing under this revision — protocol version, client info, capabilities all traveling together in a metadata block — it looks completely natural to put the tenant credential in that metadata block too. It doesn’t work, and the reason is instructive: when the client calls a tool, that call internally triggers a result-validation step, which issues its own call to list the available tools, just to check the output schema. The tool call itself accepts an application metadata argument, but the list-tools call doesn’t take one, so that internal call carries the SDK’s own protocol stamp but never your application credential — a server authorizing on that metadata block rejects its own client’s internal call, and exempting the list-tools call to work around it just reopens the hole you were trying to close.
Host: So the fix is just — don’t fight the SDK’s internal plumbing, put the credential somewhere it can’t dodge.
Guest: Exactly — authenticate on the transport, via something like the SDK’s TokenVerifier seam, so the Authorization header covers every request the transport sends, application-issued or not. Stateless protocol and per-request credentials in the body sound like the same idea, but they’re not, and that distinction generalizes to any SDK doing background refreshes or retries on your behalf, not just MCP.
11. Trade-offs and operating at scale
Host: So let’s zoom out to the operational picture. When does someone actually reach for stdio versus Streamable HTTP, given everything we just covered about credentials living at the transport layer?
Guest: stdio is the easy case — no network, no auth, the process lifecycle is the connection, so it’s great when the server runs locally alongside the host. Streamable HTTP is what you need the moment a server has to be remote or shared across multiple hosts, but you pay for that with real authentication and the routing headers stdio never has to think about.
Host: And statelessness itself — we’ve talked about the mechanics, but is it actually a win, or just a cost shifted somewhere else?
Guest: It depends entirely on what’s on the other end. For a remote server sitting behind a load balancer, it’s a clear win — horizontal scaling with no session affinity, no session store to operate or leak, trivial rolling deploys. But for a local stdio server that was never going to be load balanced in the first place, repeating protocol version and capabilities in the meta field on every single request is just pure overhead the protocol now imposes uniformly, whether you need it or not.
Host: That horizontal scaling story only covers one axis, though — what about a host that’s fanned out to many different MCP servers at once, or a shared server handling many tenants?
Guest: Right, those are separate problems statelessness doesn’t touch. A host connecting to many servers simultaneously needs the same bounded-concurrency discipline as any fan-out — semaphores on server connections, not just individual requests. And a shared remote server serving many hosts is still a multi-tenant service; removing session affinity doesn’t remove the noisy-neighbor problem, you still need per-tenant isolation. Which loops back to the many-small-servers-versus-monolith question — small focused servers keep blast radius contained, but each one you connect adds to the size of the tools list and eats more context on every model call, so past a handful of servers you need curation, not just more connections.
12. Building it yourself: labs and the bigger lesson
Host: So if someone’s absorbed all of this and wants to actually build something, where do they start?
Guest: There’s a multi-tenant MCP server lab that puts every piece we’ve discussed under test instead of just in theory — transport-level auth checked before any handler runs, per-tenant discovery that filters tools/list, and authorization checked again on tools/call because a caller can always name a tool it never saw listed. It also asserts refusals are byte-identical whether a tool is forbidden or doesn’t exist, and it proves statelessness by alternating tenants across three separate connections to one server object and checking nothing leaks.
Host: That cacheScope detail keeps nagging at me — public would tell every cache on the path that one tenant’s list is fair game for another.
Guest: Exactly, and it’s worth saying plainly: that’s a leak this protocol revision introduced the vocabulary for, not one it removed. Which is really the whole arc of this episode in miniature — statelessness didn’t make the hard problems go away, it moved them somewhere else and gave you new fields to get wrong if you’re not paying attention.
Host: So the takeaway isn’t ‘stateless is simpler,’ it’s ‘stateless relocates the state, and now you have to know exactly where it landed.’
Guest: That’s the whole skill. The session store didn’t disappear, it moved into the client’s _meta and requestState; the tenant boundary didn’t disappear, it moved into per-request checks instead of per-connection ones. Anyone can recite that MCP went stateless — the engineer worth listening to can point at each piece of state and tell you exactly where it lives now.
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Executive Summary
Section titled “Executive Summary”Before MCP, every application that wanted to connect a model to external tools and data wrote its own bespoke integration for each one — a filesystem integration for one app, a GitHub integration for another, neither reusable by anyone else. The Model Context Protocol (MCP) standardizes that connection so an integration written once, as an MCP server, works with any compliant host.
The 2026-07-28 revision changed how that connection works at the wire level, and the change is
the most instructive thing about the protocol right now: MCP removed its initialize handshake
and its protocol-level session, making every request self-describing. That is not a cosmetic
revision. A stateful protocol forces session affinity, which forces sticky routing, which is what
makes a service hard to scale horizontally — the exact lesson
Module 2 teaches about state ownership, arriving here as
a protocol redesign. This module covers MCP’s architecture, its three primitives, the security
model that split enables, and what the stateless revision changed.
Warning
This module is verified against the 2026-07-28 specification. Material written for
2025-11-25 — which is most of what currently exists online, including many SDK tutorials —
describes an initialize handshake, session IDs, and an HTTP+SSE transport that the current
revision removed or deprecated. Check which revision any MCP source targets before following it.
Mental Model
Section titled “Mental Model”MCP separates three roles, and the split is the whole design:
- Host — the application orchestrating one or more model conversations (a chat client, an IDE). It owns the user-facing experience and the decision to connect to any given server.
- Client — a connection manager living inside the host, one per connected server. It handles the protocol mechanics so the host doesn’t have to.
- Server — a separate process exposing capabilities (tools, resources, prompts) over the protocol. Servers are meant to be small and focused — one server per integration, not one server trying to expose everything.
This should feel familiar: it’s the same aggregation pattern as
labs/async-ai-gateway aggregating multiple LLM providers behind
one interface, applied to tool integrations instead of model providers. A host connects to many
independently-maintained MCP servers and presents their combined capabilities to the model as one
coherent context — the same “many small, focused things behind one aggregation point” shape as
Module 4’s gateway layer.
Architecture
Section titled “Architecture”flowchart LR
subgraph Host["Host application (e.g. an IDE or chat client)"]
C1[Client 1]
C2[Client 2]
end
C1 <-->|"JSON-RPC over stdio"| S1["Server: filesystem\n(exposes Resources)"]
C2 <-->|"JSON-RPC over Streamable HTTP"| S2["Server: GitHub\n(exposes Tools)"]
Host -->|"aggregated tools + resources + prompts"| Model[Model]The one-client-per-server design is deliberate: each connection is independently scoped (a compromised or misbehaving server affects only its own client, not the whole host) and independently transported — nothing requires every server to use the same transport, as the diagram’s stdio-and-HTTP mix shows.
What that isolation no longer implies is a stateful connection. Under 2026-07-28, a client is a
connection manager, not a session holder:
flowchart TB
subgraph Before["2025-11-25: session-bound"]
direction TB
BC[Client] -->|"initialize → session id"| BLB[Load balancer]
BLB -->|"every later request must return<br/>to the same instance"| BS1["Server instance A<br/>(holds session state)"]
BLB -.->|"cannot route here"| BS2["Server instance B"]
end
subgraph After["2026-07-28: stateless"]
direction TB
AC[Client] -->|"Mcp-Method + Mcp-Name headers<br/>_meta: version + capabilities"| ALB[Load balancer]
ALB -->|"routes on headers alone"| AS1["Server instance A"]
ALB -->|"any instance can serve it"| AS2["Server instance B"]
ALB -->|"no affinity required"| AS3["Server instance C"]
end
Before -.->|"removing the handshake removes<br/>the affinity requirement"| AfterDeep Dive
Section titled “Deep Dive”Three primitives, three different owners of the invocation decision. This is the detail most worth internalizing, because it’s a security boundary, not just a taxonomy:
- Tools are model-invokable — the model decides to call one, the same function-calling mechanics Module 5 covers, just standardized across servers.
- Resources are application-controlled data — files, database rows, search results — that the host, not the model, decides to attach to context (though some hosts also let the model request a resource read explicitly).
- Prompts are user-controlled templates — often surfaced as slash commands or menu items in a host’s UI — that the user, not the model or the application, decides to invoke.
Choosing the right primitive for a capability is itself a design decision: exposing “read this file” as a resource keeps the read under host/user control, while exposing it as a tool hands the model autonomy to decide when to read it — a real trade-off covered below.
Statelessness (SEP-2575). Every MCP message is still JSON-RPC 2.0, but there is no longer an
initialize/initialized handshake and no protocol-level session. The protocol version, client
information, and client capabilities that were once exchanged once at connection time now travel in
_meta on every request, and a client that wants server capabilities up front calls
server/discover explicitly rather than receiving them as handshake output. Every request is
self-contained.
Verified on the wire against the official Python SDK (mcp 2.3.0, and 2.0.0 before it), a modern connection issues no
initialize at all — the first method on the wire is server/discover, and the three _meta
keys carried on every subsequent request are io.modelcontextprotocol/protocolVersion,
io.modelcontextprotocol/clientInfo, and io.modelcontextprotocol/clientCapabilities. The SDK
lists 2026-07-28 as its only modern protocol version, with 2024-11-05 through 2025-11-25
retained separately as legacy handshake versions.
Routing headers (SEP-2243). Streamable HTTP now requires an Mcp-Method header on every
request, plus an Mcp-Name header on the requests that name a specific thing — tools/call,
resources/read, and prompts/get. The point is operational: a load balancer, API gateway, or
rate limiter can route and meter on the operation without parsing the JSON-RPC body. Servers must
reject requests whose headers and body disagree, which closes the obvious spoofing gap where a
request routes as one operation and executes as another.
Multi-Round-Trip Requests (SEP-2322). Without long-lived streams, a server can no longer
initiate a request back to the client mid-call. Instead, a tool that needs more input returns a
result with resultType: "input_required", an inputRequests object describing what it needs, and
an opaque requestState string. The client gathers the input and retries the call with
inputResponses and the echoed requestState. The server resumes from that state rather than
holding the original call open.
Engineering Note
requestState is the seam where “stateless protocol” meets “stateful interaction.” The protocol
stopped holding the state; it did not stop existing. The server now hands it to the client to
carry and hand back — the same continuation-passing move as a resumable upload token or a
paginated cursor. Whether that state is signed, encrypted, or expiring is a server design
decision the specification leaves open, which makes it a security decision worth being explicit
about.
Discovery and caching (SEP-2549). Discovery within each primitive keeps the same shape: list,
then invoke — tools/list then tools/call, resources/list then resources/read,
prompts/list then prompts/get. List and resource-read results now also carry ttlMs and
cacheScope, modeled on HTTP Cache-Control, so a client knows how long a tools/list response
stays fresh and whether it is safe to share that cached response across users.
Deprecations. The older HTTP+SSE transport is deprecated (SEP-2596), as are the Roots, Sampling, and Logging primitives and Dynamic Client Registration (SEP-2577, PR #2858). The lifecycle policy sets a minimum twelve-month window between deprecation and earliest removal, measured from the revision that deprecates — which puts Roots, Sampling, Logging, and DCR at the first revision released on or after 2027-07-28. HTTP+SSE is the exception: reclassified under the transition provisions, it becomes eligible three months after SEP-2596 reaches Final. All of them still work; none should carry new work, and the transport will go first.
Research Note
The canonical source for exact JSON-RPC message shapes, _meta field definitions, header rules,
and the authorization model. This module covers the architecture and the reasoning; the
specification covers the wire format. Individual changes are traceable through their SEP numbers.
Source: Model Context Protocol specification 2026-07-28, modelcontextprotocol.io
Implementation
Section titled “Implementation”The shape of a minimal MCP tool definition and a stateless request handler — illustrative of the protocol’s structure, not a copy of any specific SDK’s API surface:
A tool definition, and the header/body agreement check a stateless server owes its callers
from __future__ import annotations
from dataclasses import dataclassfrom typing import Any
@dataclassclass ToolDefinition: name: str description: str input_schema: dict[str, Any] # JSON Schema describing expected arguments
async def handle_tools_list() -> dict[str, Any]: tools = [ ToolDefinition( name="search_issues", description="Search open issues in the configured repository by keyword.", input_schema={ "type": "object", "properties": {"query": {"type": "string"}}, "required": ["query"], }, ), ] return { "tools": [t.__dict__ for t in tools], # Modeled on HTTP Cache-Control: how long this list stays fresh, and # whether it is safe to reuse across users (SEP-2549). "ttlMs": 300_000, "cacheScope": "public", }
async def handle_tools_call( name: str, arguments: dict[str, Any], headers: dict[str, str]) -> str: # The routing headers let a gateway dispatch without parsing this body, so # the server must confirm they describe the call it actually received — # otherwise a request can route as one operation and execute as another. if headers.get("Mcp-Method") != "tools/call" or headers.get("Mcp-Name") != name: raise ValueError("routing headers disagree with the request body")
if name != "search_issues": raise ValueError(f"unknown tool: {name}")
query = arguments.get("query") if not isinstance(query, str) or not query.strip(): raise ValueError("'query' must be a non-empty string")
results = await search_issues(query) # the server's own integration logic return format_results(results)Two independent validations, for two different threat models. handle_tools_call validates
arguments because the schema in tools/list is a hint to the model, not a guarantee the server
can skip checking — the model may hallucinate a malformed argument, or an adversarial client may
send one deliberately. It separately validates that the routing headers match the body, because
those headers are consumed by infrastructure that never sees the body, and a mismatch means
something upstream made a decision on false information.
Production Example
Section titled “Production Example”An IDE host connects to two MCP servers for one coding task: a filesystem server exposing the
current project as resources, and a remote GitHub server exposing tools to search issues
and open pull requests. The user attaches a specific file as context (a resource, under the
host/user’s control per this module’s Deep Dive) — the model reads its content, understands the bug
being fixed, then decides to call the GitHub server’s create_pull_request tool once it has a fix
ready.
Because the GitHub server is stateless, its operator runs it as an ordinary horizontally-scaled
HTTP service: three replicas behind a load balancer, no sticky sessions, Mcp-Method used to rate
limit tools/call more aggressively than tools/list, and a rolling deploy that drops connections
without breaking anyone’s in-flight work. Under the previous revision, every one of those
properties needed either session affinity or shared session storage.
Failure Modes
Section titled “Failure Modes”Tool poisoning via adversarial descriptions
A malicious or compromised server can write a tool description crafted to manipulate the model’s behavior beyond what the tool actually does — since tool descriptions are part of the model’s context, this is the same prompt-injection risk Module 4 covers, arriving through a server’s self-reported metadata instead of user input. Only connect to servers whose provenance you trust.
Servers trusting model-supplied arguments
A server that skips validating tools/call arguments against its own schema — assuming the model
always produces well-formed input because it was given a schema — is one hallucinated or
adversarial argument away from a bug or a security incident. Validate server-side, always, per
this module’s Implementation section.
Trusting routing headers without checking them against the body
Mcp-Method and Mcp-Name exist so infrastructure can route without parsing bodies — which
means a server that never verifies they agree with the body lets a caller route a request as a
cheap, permitted operation while executing an expensive or restricted one. The specification
requires servers to reject disagreement; a hand-rolled server is exactly where that check gets
forgotten.
Carrying 2025-era session assumptions into a stateless server
Code ported from the previous revision often keeps per-connection state — a cached capability
set, an authenticated identity, an in-progress operation — behind an implicit assumption that the
same client returns to the same instance. Under 2026-07-28 nothing guarantees that, so the
symptom is intermittent and load-dependent: fine on one replica in development, failing under a
load balancer in production.
Unbounded or forgeable requestState in Multi-Round-Trip Requests
requestState is state the server hands to the client and trusts on return. A server that does
not sign, expire, or size-bound it has created a client-controlled input that resumes
server-side execution — the same class of mistake as trusting an unsigned cookie to carry
authorization.
A crashed stdio server with no reconnect logic
Since a stdio server’s process lifecycle is its connection lifecycle, a crashed server process silently ends that connection. A host with no detection or reconnect logic for this just loses that server’s capabilities mid-session with no visible error.
Trade-offs
Section titled “Trade-offs”stdio vs. Streamable HTTP transport
stdio is simpler — no network, no authentication, the process lifecycle IS the connection — but only works for a server running locally alongside the host. Streamable HTTP supports remote and shared servers serving many hosts, at the cost of needing real authentication and the routing headers the local case never has to think about. Note that the older HTTP+SSE transport is deprecated (SEP-2596) and is no longer a third option for new work.
Stateless protocol vs. per-request overhead
Statelessness buys horizontal scaling without affinity, trivial rolling deploys, and no
session store to operate or leak. It costs bytes and repetition: protocol version, client info,
and capabilities ride along in _meta on every request rather than being negotiated once. For
a remote server behind a load balancer that trade is clearly worth it; for a local stdio server
that was never going to be load balanced, it is pure overhead the protocol now imposes uniformly.
One focused server per integration vs. one server exposing everything
Many small, focused servers (the pattern this module’s Mental Model section recommends) keep each integration’s blast radius small and let hosts connect only to what a task needs. One monolithic server exposing many unrelated capabilities is simpler to deploy as a single unit, but a compromise or bug in it affects everything it exposes at once — the same specialization trade-off Module 5 covers for agents themselves.
Exposing a capability as a tool vs. a resource
The same underlying capability — say, “read this file” — can be exposed as a resource (host/user decides when to attach it) or a tool (model decides when to call it). Resources keep a human or the host application in control of what enters context; tools give the model autonomy to fetch what it decides it needs, at the cost of less predictable context growth and a wider action surface for the model to misuse.
Security
Section titled “Security”- Tool and resource descriptions are part of the model’s context — treat an untrusted server’s self-reported metadata with the same suspicion as any other untrusted input, per the tool poisoning failure mode above.
- Validate every
tools/callargument server-side, regardless of the schema the model was given — the schema constrains what a well-behaved client sends, not what a server should assume it received. - Reject requests whose routing headers disagree with their body, since infrastructure upstream has already made routing, rate-limiting, and authorization decisions on those headers alone.
- Treat
requestStateas untrusted on return. Sign it, bound its size, and expire it; it is client-held state that resumes server-side execution. The official SDK does exactly this — it ships an AES-GCMrequestStatecodec that binds the state to the authenticated principal, which is a good indication of the threat model the specification authors had in mind. - Put credentials in the transport, not in
_meta. “Every request is self-describing” makes per-request application credentials look natural, and_metalooks like the obvious place for them. It is the wrong place: an SDK makes protocol calls of its own that application code never issues, and those carry no application metadata. Concretely, on the Python SDK acall_tool()for a tool the client has not listed yet produces two server-visible requests — thetools/callitself, and an internaltools/listissued byvalidate_tool_result()to fetch the output schema. Both carry the protocol_metastamp (protocolVersion,clientInfo,clientCapabilities), because the SDK adds that to every outgoing request. Only the first carries yours: the second is sent by the SDK itself, so an application credential cannot reach it — even thoughlist_tools()acceptsmeta=when your code is the caller. A server authorizing on a token in_metatherefore rejects its own client, and a server that quietly allows unauthenticatedtools/listto compensate has opened exactly the hole it was trying to close. Authorization belongs on the Authorization header, where every request the transport sends — application-issued or not — carries it. - Follow the hardened authorization model.
2026-07-28requires RFC 9207 issuer validation and binds credentials to the issuing authorization server. Dynamic Client Registration is deprecated in favour of Client ID Metadata Documents (CIMD) — new integrations should use CIMD. - Require host-level consent for sensitive operations — write actions, destructive operations, anything with real-world cost — as a non-optional gate, not a configuration flag defaulting to off.
- Scope each server connection to least privilege — a host shouldn’t connect to a server with broader capability (filesystem access, API scope) than the current task actually needs.
Performance
Section titled “Performance”_metaoverhead is now per-request, not per-connection. Protocol version, client info, and capabilities repeat on every call. It is small in absolute terms, but it is a real cost that scales with request count rather than connection count — worth measuring on a chatty client before assuming it is free.ttlMsandcacheScopeturn discovery into a cacheable operation. A client that honours them stops re-fetchingtools/liston every turn, and acacheScopemarking a response shareable lets a gateway serve many users from one cached copy. Ignoring these fields silently gives up the main performance win of the revision.tools/listresponse size grows with every connected server, and that whole list typically enters the model’s context on each call — the same context-growth cost Module 5 covers for agent loops, driven here by how many servers a host has connected to rather than how many loop iterations have run.- stdio subprocess startup cost matters if servers are spawned fresh per session rather than reused across a host’s lifetime — the same connection-reuse logic Module 3 covers for network connections applies to process spawning here.
- Resource read latency depends entirely on what’s behind the resource — a slow API-backed resource server blocks exactly like any other slow dependency, and needs the same timeout and bounded-concurrency discipline from Module 1.
Scaling
Section titled “Scaling”- Statelessness is what makes a remote MCP server ordinary to scale. With no session to pin a
client to an instance, a server scales horizontally behind a plain load balancer, tolerates
rolling deploys without draining sessions, and needs no shared session store. This is the single
biggest operational consequence of
2026-07-28, and it is why the revision matters more to platform teams than to individual tool authors. - Header-based routing makes per-operation policy possible at the edge. Because
Mcp-Methodis visible without body parsing, a gateway can rate limittools/calldifferently fromtools/list, route expensive operations to dedicated capacity, or shed load by operation class — the same edge-policy pattern Module 4 applies to model traffic. - A shared, remote MCP server serving many hosts is still a multi-tenant service and needs the same per-tenant isolation Module 2 covers for any shared backend — statelessness removes session affinity, not the noisy-neighbour problem.
- A host connecting to many MCP servers simultaneously needs the same bounded-concurrency thinking as any fan-out of concurrent connections — see Module 1’s treatment of semaphores, applied to server connections instead of individual requests.
- Tool and resource list growth across many connected servers eventually needs curation — not every capability from every connected server belongs in every model call’s context; selective or dynamic exposure becomes necessary well before a host is connected to more than a handful of servers.
Interview Questions
Section titled “Interview Questions”Explain the host/client/server split in MCP and why it exists.
The host orchestrates the overall experience and decides what to connect to; a client is a connection manager per server, isolating one server’s connection from another’s; the server exposes capabilities over a standard protocol so it can be reused across any compliant host. The split keeps a misbehaving or compromised server’s blast radius limited to its own client connection.
What's the difference between a tool, a resource, and a prompt in MCP?
They differ in who decides to invoke them: the model decides for tools, the host or user decides for resources, and the user explicitly decides for prompts (often via a UI affordance like a slash command). The right primitive for a capability depends on who should control that decision.
MCP removed its initialize handshake and session in 2026-07-28. What does that change operationally?
It removes session affinity as a deployment requirement. A remote MCP server becomes an ordinary
stateless HTTP service: horizontally scalable behind a plain load balancer, safe to roll deploys
through, with no session store to operate. The cost is that protocol version, client info, and
capabilities now ride in _meta on every request instead of being negotiated once, and any
server code that implicitly assumed per-connection state has to be rewritten to stop assuming it.
What a strong answer connects it to
The best answers do not treat this as MCP trivia — they name it as the standard stateful-to-
stateless service migration, the same reasoning behind stateless application servers and
token-based sessions, and note that the state did not vanish so much as move: to the client, in
_meta and requestState. Naming where the state went, and what now has to secure it, is the
distinguishing move.
How would you secure a host connecting to third-party MCP servers?
Treat every server’s tool/resource descriptions as untrusted input (tool poisoning), require
human approval gates before sensitive tool invocations, scope each connection to least privilege,
and validate that servers you don’t control are actually validating their own inputs before
trusting their behavior at all. On the authorization side, 2026-07-28 requires RFC 9207 issuer
validation and binds credentials to the issuing authorization server, with CIMD replacing
deprecated Dynamic Client Registration.
Why do Mcp-Method and Mcp-Name exist, and what must a server do about them?
They let load balancers, gateways, and rate limiters route and meter on the operation without parsing the JSON-RPC body — cheaper, and it works for infrastructure that should not be reading request bodies at all. The obligation they create is that a server must reject any request whose headers and body disagree; otherwise a caller can have a request routed and rate-limited as one operation while it executes as another.
When would you choose stdio vs. Streamable HTTP transport for an MCP server?
stdio for a server running locally alongside its host, with no need for remote access or multiple simultaneous clients — simpler, no auth needed. Streamable HTTP when the server needs to be remote, shared across multiple hosts, or accessed over a network at all. HTTP+SSE is deprecated and is not a choice for new work.
Hands-on Lab
Section titled “Hands-on Lab”Multi-Tenant MCP Server is a real MCP server on the
official SDK, speaking this revision of the protocol over Streamable HTTP: transport-level tenant
authentication, per-tenant discovery, refusals that leak nothing, cacheScope: private on
tenant-scoped listings, and statelessness proven by test rather than asserted.
Warning
A lab that implements JSON-RPC with MCP-shaped method names is a protocol-learning exercise,
not an MCP implementation, and should say so. Calling something MCP-compliant means it
interoperates with real hosts — which under 2026-07-28 means self-describing requests,
Mcp-Method/Mcp-Name handling, and the current authorization model. Build against an official
SDK, or enumerate exactly which parts of the specification the exercise does and does not
implement.
Wrap a gateway capability as an MCP tool
Expose labs/async-ai-gateway’s /v1/generate endpoint as an
MCP tool server with a single tool (generate_completion), validating its input schema
server-side per this module’s Implementation section. Use an official MCP SDK so the transport
and _meta handling are specification-correct, then confirm the tool is discoverable via
tools/list and callable via tools/call from a real host.
Prove your server is actually stateless
Run two replicas of the server from the previous exercise behind a load balancer with no session affinity, and drive a multi-step interaction — including one Multi-Round-Trip Request — through it while round-robining every single request between replicas. Any failure is per-connection state you did not know you had. Then kill one replica mid-interaction and confirm the interaction still completes.
References
Section titled “References”- Model Context Protocol specification — canonical reference
for message shapes, transports,
_metafields, and the authorization model. - The 2026-07-28 specification release — the changes this module is verified against, with SEP references.
- Relevant SEPs: SEP-2575 (stateless protocol,
_meta,server/discover), SEP-2243 (Mcp-Method/Mcp-Namerouting headers), SEP-2322 (Multi-Round-Trip Requests), SEP-2549 (ttlMs/cacheScopecaching), SEP-2577 (Roots/Sampling/Logging deprecation), SEP-2596 (HTTP+SSE transport deprecation). - RFC 9207 — OAuth 2.0 authorization server issuer identification, required by the current authorization model.
- Anthropic, “Introducing the Model Context Protocol”
- Module 5: Agent Engineering — the tool-calling mechanics MCP standardizes across servers.
Revision History
Section titled “Revision History”| Version | Date | Change |
|---|---|---|
| 2.3.0 | 2026-10-08 | Re-verified against mcp 2.3.0 on the wire; corrected when the SDK’s internal tools/list happens and why the caller’s _meta misses it. |
| 2.2.0 | 2026-08-12 | Re-verified against the published specification text; fixed a cacheScope value, split the deprecation timeline, and corrected the _meta security note. |
| 2.1.0 | 2026-08-08 | Verified against the official Python SDK (mcp 2.0.0); added the transport-versus-_meta credential-placement lesson found while building against it. |
| 2.0.0 | 2026-08-08 | Rewritten against the 2026-07-28 specification: stateless protocol, routing headers, Multi-Round-Trip Requests, caching fields, and the HTTP+SSE / Roots / Sampling / Logging deprecations. |
| 1.0.0 | 2026-08-05 | Initial publication. |