Skip to content

Module 3: Networking

Listen to this page17:38
Read the transcript

1. “It’s Just HTTP” Undersells Streaming

Host: So we say ‘it’s just HTTP’ like that settles everything, but with AI infrastructure that phrase is doing a lot of hiding. These responses aren’t arriving as one neat package anymore, they’re streaming token by token, and what happens to that stream between your server and the user’s screen depends on decisions most teams never think about.

Guest: Right, and the frustrating part is the stream can look perfect on your server and still arrive as one silent block on the client, because some proxy three hops away decided to buffer it. This module is really about that whole path — connection setup overhead, which streaming protocol you actually chose, how load balancing treats long-lived connections, and where your latency budget really goes. Get any of those wrong and ‘it’s just HTTP’ turns into a very expensive lie.

2. Two Mental Models: Setup Cost vs. Work Cost, and Time-to-First-Token

Host: So let’s ground this in the mental model before we get into protocols. When you say setup cost versus work cost, what’s actually happening before any of my request even gets processed?

Guest: Every connection pays a toll before it does any real work — DNS resolution to find the address, a TCP handshake to open the connection, then a TLS handshake on top if it’s HTTPS. For a short request, that toll can be more expensive than the actual work, which is why keep-alive is the highest-leverage fix here — you pay the toll once and amortize it across every subsequent request instead of paying it again and again.

Host: Okay, but streaming adds a second wrinkle on top of that, right — it’s not just about total time?

Guest: Exactly, and this is the model people miss: for streaming, time-to-first-token matters more than total duration. An 8-second response that starts appearing in 200 milliseconds feels fast because the user watches it build; a 3-second response that arrives as one silent block feels slow even though it’s objectively quicker — and that’s entirely a perception problem created by where the bytes actually show up in time.

3. The Architecture: A Gateway-to-Client Path That Must Stay Transparent

Host: So if time-to-first-token is the perception game, what’s actually the architecture that lets those early bytes show up fast? Walk me through the path from gateway to client.

Guest: Picture the request going from gateway to client, and the rule is simple to state but easy to violate: every single hop has to forward bytes as they arrive, not wait to assemble the full response. It’s an infrastructure-wide property — one buffering proxy anywhere in that chain, and every request through it silently reverts to block delivery, even if your app code streams perfectly. The second commitment is disconnect detection has to happen at the gateway per chunk, not once at connection start, because if a client bails mid-stream, you need the gateway to notice immediately and stop pulling tokens upstream — otherwise you’re generating and paying for a response nobody’s there to receive.

4. Paying the Setup Tax: DNS, TCP, TLS, and the HTTP/1.1 → 2 → 3 Story

Host: Okay, before we even get to streaming behavior, there’s a tax you pay just to open the connection in the first place. Walk me through what that actually costs in round trips.

Guest: Every fresh HTTPS connection pays DNS resolution if it’s not cached, a TCP handshake that’s one round trip, and a TLS handshake — one round trip on TLS 1.3, or effectively zero extra with 0-RTT resumption if you’ve talked to that server before. Add it up and you’re looking at two to three round trips of pure setup before a single byte of your actual request moves, which for a server one network hop away is real time. That’s exactly why a service hammering the same upstream repeatedly should be reusing connections, not paying that tax fresh on every call — and it’s also why DNS caching is a trade-off, not a free lunch, since caching too aggressively past the TTL means you can keep routing to a dead backend for a scary long time after a failover.

Host: So once the connection’s open, that’s where HTTP/1.1 versus 2 versus 3 comes in — what actually changes for a bunch of concurrent streaming completions sharing infrastructure?

Guest: HTTP/1.1 only sends one request at a time per connection, so clients just open multiple connections to fake concurrency. HTTP/2 fixes that by multiplexing many requests over a single TCP connection — but because it’s still TCP underneath, one lost packet stalls every multiplexed stream on that connection, so streams that have nothing to do with each other end up blocking on each other’s retransmission. HTTP/3 runs on QUIC over UDP instead, giving each stream independent loss recovery, so a lost packet on one completion’s stream doesn’t stall the twenty other completions sharing that connection — which matters a lot the moment you’re multiplexing many concurrent streaming responses through the same pipe.

5. DNS Caching’s TTL Trade-off — and When It Bites

Host: So DNS caching feels like it should be a solved, boring problem, but you’re telling me it has a real trade-off built into it. What’s the tension?

Guest: Cache the resolved IP aggressively and you save resolution round trips on every request, which is great for latency. But cache it too aggressively and you’re slow to notice when infrastructure actually changes underneath you — say a load balancer fails over to a new IP. I’ve seen the exact production bug this causes: someone switches DNS, and twenty minutes later they’re still getting reports of traffic hitting the old, dead load balancer, because some fraction of clients cached the old IP past its TTL and just never re-resolved.

6. Choosing a Streaming Protocol: SSE, WebSockets, gRPC

Host: Okay, so DNS is sorted, the handshake tax is paid — now you actually have to pick how the tokens travel from your server to the browser. There’s SSE, there’s WebSockets, there’s gRPC streaming. How do you choose?

Guest: For a one-directional LLM completion — server pushing tokens to a client that isn’t sending anything back mid-stream — SSE is almost always the right call, and the reason is boring in the best way: it’s just plain HTTP with a text/event-stream content type. It works through basically every proxy and load balancer you already have without special configuration, and browsers even reconnect automatically if the connection drops. WebSockets give you full bidirectional communication, but you’re not using that for a completion stream, and you’re paying for it in infrastructure that has to explicitly support the protocol upgrade — some proxies and load balancers need extra config just to let that handshake through. gRPC streaming is efficient and strongly typed, which sounds great, but it requires HTTP/2 end-to-end, meaning every single intermediary between you and the client has to support it, and that’s a much bigger operational ask than just serving an HTTP response.

Host: So it’s really a case of matching the protocol’s capabilities to what you actually need, not what sounds most sophisticated.

Guest: Exactly — bidirectional and strongly-typed are both genuinely useful properties, just not for this shape of problem. The moment you need the client streaming data back concurrently, WebSockets earn their keep. But for the common case of a prompt in, tokens out, SSE demands the least from the path in between, and ‘demands the least from the path in between’ is exactly what you want when that path includes proxies and load balancers you don’t control.

7. Load Balancing, Latency Budgets, and Where LBs Hit Their Ceiling

Host: Okay, so once you’ve picked SSE, you still have to get the request through a load balancer without wrecking that latency budget. Walk me through the L4 versus L7 choice here.

Guest: L4 just looks at IP and port and shoves packets along — it’s fast and cheap but blind to what’s actually in the request, so it can’t route a request by tenant or API version. L7 terminates more of the connection and actually reads headers or paths, which lets you do that smarter routing, but that inspection costs real latency and CPU per request. And that cost matters more than it sounds like — a same-datacenter round trip is something like half a millisecond, but cross-continent is 100 to 150 milliseconds, so any latency an L7 hop adds is competing against a budget where the floor is already set by physics, not by your code.

Host: So what happens when the load balancer itself becomes the bottleneck, not just an added hop?

Guest: Every load balancer has its own ceiling on connections and throughput, and once you’re near it, adding more backend capacity behind it doesn’t help — the LB itself is the constraint. At that point you either scale the load balancer horizontally, or you use something like Direct Server Return, where the LB handles the incoming request but the response goes straight from the backend to the client, taking the LB out of the return path entirely.

8. Inside the Disconnect-Aware Streaming Loop

Host: So let’s actually look at the code that lives at the bottom of all this theory. There’s this is_disconnected check inside the streaming loop — walk me through why it’s called on every single chunk instead of just once when the connection opens.

Guest: Because a client can vanish at any point in a long stream, not just at the start. If you only checked once, you’d catch the person who closes the tab immediately, but you’d completely miss someone who bails halfway through a long generation. So the check sits inside the async for loop, right before you yield each chunk, and the moment it comes back true you break and stop pulling from the provider entirely.

Host: And separately from that disconnect logic, you’re also timing first token specifically, not just the whole stream duration. Why does that deserve its own percentile bucket instead of just folding into total latency?

Guest: Because a slow-starting stream and a long-running stream look identical if you only track total duration, but they need entirely different fixes. Tracking first-token time as its own p50, p95, p99 distribution lets a dashboard tell those two situations apart at a glance, rather than burying one inside the other.

9. The Real Cost of Ghost Streams

Host: Okay, let’s make this painfully concrete, because ‘latency budgets’ can feel abstract until you attach a dollar figure to it. Walk me through what actually happens when someone closes a chat tab three seconds into a ten-second response.

Guest: Without disconnect detection, nothing stops. The gateway keeps calling the upstream provider for the remaining seven seconds, and the provider keeps generating and billing tokens, because as far as it knows there’s still a client waiting. You’re paying full metered compute for a response that vanished into a closed socket with zero eyeballs on it.

Host: And that’s not a one-off annoyance, that’s a scaling problem — thousands of concurrent streams, some steady percentage of tab-closers, it just compounds. So is that why disconnected gets its own counter in StreamTelemetry instead of just folding into completed?

Guest: Exactly — so the cost is bounded to those three seconds instead of ten. But the counter matters beyond savings: a rising disconnect rate is a signal. Maybe responses are too slow and people are bailing, maybe a client has a bug that’s closing streams early. Bury that inside ‘total requests served’ and you’d never see it coming.

10. Failure Modes: Buffering Proxies, Head-of-Line Blocking, Idle Timeouts

Host: Okay, so let’s talk about how streaming actually dies in production, because I feel like this is the part that bites people who did everything right in their own code. What’s the number one culprit?

Guest: It’s almost always a buffering proxy sitting in front of the gateway. Nginx’s default is to buffer the whole response, meaning it collects the entire upstream response before forwarding a single byte — so your beautiful token-by-token stream turns into one silent block that lands all at once. And the cruel part is it’s invisible locally, because in dev your client usually talks straight to the gateway with nothing in between.

Host: So you can ship something that streams perfectly on your laptop and then watch it go silent-then-dump the moment it’s behind a real proxy. Is that the only place things quietly break, or are there other traps like this?

Guest: A few more worth knowing. HTTP/2 multiplexes several streams over one TCP connection, so one lost packet on an unrelated slow stream can stall every other stream sharing that connection — completely unrelated requests start stuttering together. There’s also idle-timeout mismatches between a proxy and a client, which cause random-looking connection resets that actually line up exactly with the timeout window once you check. Best part is you can reproduce the buffering issue in ten minutes: put nginx in front of the gateway’s stream endpoint with response buffering switched on, watch it arrive as one block, then switch buffering off and tell the proxy not to accelerate-buffer the response, and watch streaming come back instantly.

11. Securing and Scaling the Streaming Path

Host: Let’s shift to hardening this path, because a streaming endpoint that’s fast but insecure isn’t much of a win. Where does security actually bite here that’s specific to streaming, versus generic API security advice?

Guest: TLS needs to be everywhere, including internal service-to-service hops, not just the public edge, and someone needs to actively watch certificate expiry because that’s a boring, entirely preventable cause of outages that still takes down production systems constantly. For zero-trust designs, mTLS on service-to-service traffic establishes identity cryptographically per-connection instead of assuming trust from network location. DNS itself is a trust boundary, not just a lookup table, so spoofing and cache poisoning are real risks, and DNSSEC belongs in the picture where the threat model justifies it. And a sneaky one: a permissive Access-Control-Allow-Origin wildcard on a streaming endpoint that also handles authenticated requests is a common, easy-to-miss exposure — layer rate limiting at both the edge and the application so one bug in one layer isn’t the only thing standing between you and abuse.

Host: Okay, and once that’s locked down, what breaks first when you actually try to scale this thing across regions and high fan-out?

Guest: Connection pool sizing becomes a real constraint — too small and requests queue waiting for a connection, too large and you risk exhausting file descriptors, or behind NAT, ephemeral ports, and SNAT port exhaustion is a genuinely painful ceiling to hit in NAT-heavy cloud topologies at scale. And for multi-region deployments, you need latency-aware routing like geo-DNS or anycast, because without it a user gets routed to a healthy backend in the wrong region, trading availability for a latency budget far worse than what those setup-cost numbers from earlier suggested you’d need to accept.

12. Closing: Interview-Ready Takeaways and Where to Go Deeper

Host: So if someone drops us into an interview tomorrow and asks us to name the signals that show real networking judgment for streaming AI systems, what’s the short list? Give us the version that sounds sharp in five sentences.

Guest: Three things, and we’ve already built each one out in detail earlier, so here they are as the compressed, interview-ready version. First, the latency breakdown — DNS, TCP handshake, TLS handshake, request transmission, server processing, first byte — treated as distinct, separately measurable stages, the way we covered under the setup tax and latency budget discussions. Second, the SSE-over-WebSockets call for a one-directional completion stream, for the exact reasons we laid out when choosing a streaming protocol: plain HTTP, no upgrade handshake, no unused bidirectional overhead. And third, the disconnect answer we built out in the disconnect-aware streaming loop and the ghost-streams discussion: check state on every chunk, stop pulling from upstream immediately, tie it to its own counter, and remember that in AI infrastructure an undetected disconnect is metered per-token spend on a response nobody’s watching. If you want to go deeper than the interview answer, Grigorik’s High Performance Browser Networking is still the best treatment of the TCP and TLS mechanics, the HTTP/2 and HTTP/3 RFCs cover the multiplexing and head-of-line-blocking details directly, and the async-ai-gateway lab is the actual disconnect-aware implementation this whole module was drawn from — that’s where the mental models turn into working code.

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

“It’s just HTTP” undersells what actually happens between a client request and a response — especially for AI infrastructure, where responses increasingly stream token-by-token instead of arriving all at once. This module covers the network layer decisions that determine whether that streaming response reaches a user smoothly or gets silently buffered into one big chunk by a proxy three hops away: connection lifecycle overhead, streaming protocol choice, load balancing, and the latency budget every request actually spends.

Every network request pays a setup cost before it pays a work cost, and for short requests the setup cost often dominates: DNS resolution, then a TCP handshake, then (for HTTPS) a TLS handshake, before a single byte of the actual request is sent. Connection reuse (keep-alive) is the single highest-leverage optimization in this module because it amortizes that setup cost across many requests instead of paying it every time.

For streaming responses specifically — the shape of every LLM completion API — add a second mental model on top: time-to-first-byte (or first-token) is a separate, often more important, metric than total duration. A response that takes 8 seconds total but starts streaming within 200ms feels fast to a user watching text appear. A response that takes 3 seconds total but arrives as one block after a 3-second silence feels slow — even though it’s objectively faster end to end.

Streaming is a UX decision as much as a protocol decision

The reason AI infrastructure cares about this module more than most backend systems: token-by-token streaming is now the default UX for LLM output, and every intermediary between your gateway and the user — proxies, CDNs, load balancers — has to be streaming-transparent for that UX to survive the trip. One misconfigured buffering proxy and a beautifully streamed response arrives at the browser as one silent block.

The path a streamed LLM response takes from gateway to client, and the two places production systems most often get it wrong:

Two architectural decisions this diagram makes explicit:

  1. Every hop between the gateway and the client has to forward bytes as they arrive, not buffer the full response first. This is an infrastructure-wide property, not something the application layer can fix alone — a single buffering proxy anywhere in the path defeats streaming for every request that passes through it.
  2. Disconnect detection happens at the gateway, per chunk, not just once at the start. A client that closes its connection mid-stream should stop the gateway from continuing to pull (and pay for) tokens from the upstream provider — covered concretely in the Production Example below.

Connection setup cost. A fresh HTTPS connection pays, in order: DNS resolution (unless cached), a TCP handshake (one round trip), and a TLS handshake (one round trip for TLS 1.3, effectively zero extra round trips with TLS 1.3’s 0-RTT resumption for a previously-seen server). For a request to a server one network hop away, that’s easily 2-3 round trips of pure setup before any application data moves — which is why a service making many short-lived requests to the same upstream should almost always reuse connections instead of opening a new one per request.

HTTP/1.1 vs HTTP/2 vs HTTP/3. HTTP/1.1 sends one request at a time per connection (pipelining existed but was never widely deployed), so clients open multiple connections to get concurrency. HTTP/2 multiplexes many requests over one TCP connection — but because it’s still built on TCP, a single lost packet stalls every multiplexed stream on that connection until it’s retransmitted (TCP head-of-line blocking, now shared across requests that have nothing to do with each other). HTTP/3, built on QUIC over UDP, gives each stream independent loss recovery — a lost packet on one stream no longer stalls the others. For AI infrastructure specifically, this matters most when multiplexing many concurrent streaming completions over a shared connection.

DNS resolution and caching. DNS TTLs create a real trade-off: cache aggressively and save resolution round trips, or cache conservatively and pick up infrastructure changes (a failed-over load balancer IP, a rotated upstream) quickly. A client that caches a resolved IP far past its TTL can keep routing to a dead backend for uncomfortably long after a failover — this is a common, confusing production bug (“why is traffic still hitting the old load balancer 20 minutes after we switched DNS?”).

Streaming protocol choice: SSE vs. WebSockets vs. gRPC streaming. Server-Sent Events (SSE) is the simplest option for one-directional, server-to-client token streaming — it’s plain HTTP (text/event-stream), works through most existing HTTP infrastructure without special handling, and reconnects automatically in browsers. WebSockets give full bidirectional communication, at the cost of more infrastructure that needs to explicitly support the protocol upgrade (some proxies and load balancers need extra configuration). gRPC streaming is efficient and strongly typed, but requires HTTP/2 support end-to-end, including through any intermediary — a much bigger operational ask than “just serve text/event-stream.” For a one-directional LLM completion stream, SSE is usually the right default specifically because of how little infrastructure it demands.

Load balancing: L4 vs. L7. A Layer 4 (transport-layer) load balancer routes based on IP and port without looking at application content — fast, low-overhead, but blind to what’s actually in the request. A Layer 7 (application-layer) load balancer can route on HTTP headers, paths, or even request content — enabling things like routing by tenant or API version — at the cost of terminating and inspecting more of the connection, adding latency and CPU cost per request.

Research Note

A round trip within the same datacenter is roughly 0.5ms; a round trip to a different continent is 100-150ms. These numbers are decades old and the constants have shifted somewhat, but the ratios between them — same-rack vs. same-datacenter vs. cross-region — still shape every latency budget decision in this module.

Source: Jeff Dean, "Numbers Every Programmer Should Know"

The disconnect-aware SSE streaming loop, from the actual production endpoint in labs/async-ai-gateway:

Client-disconnect-aware SSE streaming, from production_app.py

@app.post("/v1/stream")
async def stream(request: Request, request_body: GenerateRequest, ...) -> StreamingResponse:
async def event_stream() -> AsyncIterator[str]:
log_event("stream.started", tenant_id=tenant, provider=request_body.provider)
try:
async for chunk in gateway.stream(request_body.prompt, provider_name=request_body.provider):
if await request.is_disconnected():
log_event("stream.client_disconnected", tenant_id=tenant)
break
yield format_sse_chunk(chunk)
finally:
log_event("stream.finished", tenant_id=tenant)
return StreamingResponse(event_stream(), media_type="text/event-stream")

The is_disconnected() check runs on every chunk, not once at the start — a client can disconnect at any point during a long stream, and checking only once would miss every disconnect after the first token. This is the direct implementation of the Architecture section’s second design decision: stop pulling tokens from the upstream provider the moment nobody is listening.

Pairing that with time-to-first-token measurement — a metric this module’s Mental Model section argues matters more than total duration — from stream_telemetry.py:

def record_first_token(self, started_at: float) -> None:
self.first_token_ms.append((time.perf_counter() - started_at) * 1000)

Tracked as its own percentile distribution (p50/p95/p99), separate from total stream duration — so a dashboard can distinguish “streams are starting slowly” from “streams are running long,” which need entirely different fixes.

A user opens a chat UI, sends a prompt, and closes the browser tab three seconds into a ten-second LLM response. Without disconnect detection, the gateway keeps pulling tokens from the upstream provider for the full ten seconds — burning real, metered compute cost on a response nobody will ever see. With the is_disconnected() check from the Implementation section, the gateway notices the closed connection on the very next chunk and stops pulling from the provider immediately.

At meaningful scale — thousands of concurrent streams, some fraction of users always closing tabs mid-response — this isn’t a cleanliness nicety, it’s a direct, measurable cost line. This is also why StreamTelemetry tracks disconnected as its own counter, separate from completed: a rising disconnect rate is a real production signal (are responses too slow to keep users waiting for the end? is a client bug closing connections early?) that a “total requests served” metric would completely hide.

Buffering proxies breaking streaming

A reverse proxy configured with response buffering enabled (a common default, e.g. nginx’s proxy_buffering on) collects the entire upstream response before forwarding any of it — turning a beautifully streamed, token-by-token response into one silent block that arrives all at once. This is the single most common way streaming quietly breaks in production, and it’s invisible in local development where the client often talks directly to the gateway with no proxy in between.

HTTP/2 head-of-line blocking across multiplexed streams

Multiple concurrent streaming completions sharing one HTTP/2 connection all stall together if a single TCP packet is lost — a completely unrelated slow or congested stream can visibly stutter every other stream sharing that connection, a direct consequence of HTTP/2 being multiplexed over one TCP connection (see Deep Dive).

Overly aggressive DNS caching

A client or resolver caching a resolved IP well past its TTL keeps routing to infrastructure that no longer exists after a failover or migration — a confusing failure because it looks like “the new backend isn’t working” when the real problem is that some fraction of traffic never re-resolved.

Idle connection timeouts closing keep-alive connections mid-use

A load balancer or proxy with an idle timeout shorter than the client’s keep-alive assumption closes connections the client believes are still open — producing intermittent “connection reset” errors that look random but correlate exactly with the idle-timeout window whenever you check.

The cold-connection tax on short requests

A service making many short, one-off requests to a downstream that doesn’t reuse connections pays the full DNS + TCP + TLS setup cost on every single request — often more expensive than the request’s actual work. If p50 latency looks suspiciously close to “connection setup time” for a simple request, this is the first thing to check.

SSE vs. WebSockets vs. gRPC streaming

SSE demands the least from surrounding infrastructure (plain HTTP, works through most existing proxies) but is one-directional only. WebSockets add bidirectional communication at the cost of explicit infrastructure support for the protocol upgrade. gRPC streaming is efficient and typed but requires HTTP/2 end-to-end, including every intermediary — the biggest operational ask of the three. Per the Deep Dive section, SSE is the right default for one-directional LLM token streaming specifically because of how little it demands from the path in between.

L4 vs. L7 load balancing

L4 is fast and simple but can’t route on request content — no per-tenant or per-API-version routing. L7 enables that richer routing at the cost of terminating and inspecting more of the connection per request, which is real added latency and CPU at scale. Many production topologies use both: an L4 balancer at the edge for raw throughput, L7 routing behind it where the smarter decisions actually need to happen.

HTTP/2 multiplexing vs. HTTP/1.1 simplicity

HTTP/2’s multiplexing removes the “open many connections for concurrency” workaround HTTP/1.1 clients rely on, at the cost of the shared-connection head-of-line blocking covered in Failure Modes. For a service running many concurrent long-lived streams, that trade-off deserves a real look rather than an assumption that “HTTP/2 is strictly better.”

  • TLS everywhere, including internal service-to-service traffic, not just the public edge — and active monitoring for certificate expiry, which is a boring, entirely preventable cause of production outages.
  • mTLS for service-to-service communication in zero-trust network designs, so identity is established cryptographically per-connection rather than assumed from network location.
  • DNS spoofing and cache poisoning risk — DNSSEC where the threat model justifies it, and awareness that DNS is a trust boundary, not just a lookup table.
  • CORS misconfiguration on streaming endpoints — a permissive Access-Control-Allow-Origin: * on an endpoint that also handles authenticated requests is a common, easy-to-miss exposure.
  • Rate limiting at both the edge and the application layer — defense in depth, so an application-layer bug in rate limiting isn’t the only thing standing between the service and abuse.
  • Connection reuse is the highest-leverage latency fix for chatty services — per the Mental Model section, it removes the DNS/TCP/TLS setup cost from every request after the first.
  • Time-to-first-byte (or first-token) as a first-class metric, tracked separately from total duration — exactly what StreamTelemetry’s percentile breakdown does, and exactly the distinction the Mental Model section opens with.
  • Latency budgets should be decomposed by hop — DNS, TCP/TLS setup, time in a load balancer’s queue, upstream processing, and transfer time are each worth their own number; a single end-to-end latency figure can’t tell you which hop to fix.
  • Connection pool sizing becomes a real constraint at high fan-out — too small and requests queue waiting for a connection; too large and you risk exhausting file descriptors or, behind NAT, ephemeral ports (SNAT port exhaustion is a genuinely painful, easy-to-hit ceiling in NAT-heavy cloud network topologies at scale).
  • Load balancers have their own connection and throughput ceilings — scaling past them usually means horizontal LB scaling or techniques like Direct Server Return (DSR) that take the load balancer out of the response path.
  • Multi-region traffic needs latency-aware routing (geo-DNS, anycast) — without it, a user in one region can be routed to a healthy backend in a different region, trading availability for a much worse latency budget than the Mental Model section’s numbers would suggest is necessary.

Walk me through what happens between a client sending a request and receiving the first byte of response.

DNS resolution (or cache hit), TCP handshake, TLS handshake (unless 0-RTT resumption applies), request transmission, server processing, and the first response byte — each a distinct, separately measurable stage per this module’s Deep Dive and Performance sections.

Why choose SSE over WebSockets for streaming LLM output?

The completion stream is one-directional (server to client) and SSE is plain HTTP that works through existing infrastructure with no special upgrade handling — WebSockets’ bidirectional capability is unused overhead for this specific use case, and it demands more from every intermediary in the path.

How do you detect and handle a client disconnecting mid-stream?

Check the connection’s disconnected state on every chunk of the stream — not just once at the start — and stop pulling from the upstream provider the moment it’s detected, per this module’s Implementation section. Tie it to a metric (a disconnect counter, separate from completions) so a rising disconnect rate is visible as its own signal.

Why the cost angle matters here specifically

For AI infrastructure specifically, undetected disconnects aren’t just wasted server cycles — they’re wasted, metered, per-token upstream provider cost on a response nobody will ever see. That’s the detail that separates a generic “networking best practice” answer from one that shows AI-infrastructure-specific judgment.

What's HTTP/2 head-of-line blocking, and how does HTTP/3 address it?

HTTP/2 multiplexes many streams over one TCP connection, so a single lost packet stalls every stream on that connection until TCP retransmits it — even streams with no relation to the lost packet. HTTP/3, built on QUIC over UDP, gives each stream independent loss recovery, so a lost packet only stalls the stream it belonged to.

Hands-on Lab

The gateway’s /v1/stream endpoint implements exactly the disconnect-aware SSE pattern this module covers, with time-to-first-token and disconnect-rate telemetry already wired up. Read the full lab documentation →

labs/async-ai-gatewayproduction-ready

Reproduce the buffering-proxy failure mode locally

Put a local nginx instance in front of the gateway’s /v1/stream endpoint with proxy_buffering on (the default) and observe the response arrive as one block instead of streaming. Then set proxy_buffering off and X-Accel-Buffering: no and confirm streaming behavior returns. This is the single most common way this module’s central failure mode shows up in real deployments — reproducing it locally makes it recognizable instantly in production.

Before shipping a streaming endpoint to production

  • Every intermediary in the path (proxies, CDN, load balancer) is configured to disable response buffering for this endpoint - The server checks for client disconnect on every chunk, not just at stream start - Time-to-first-byte/token is tracked as a metric separate from total stream duration
  • A disconnect-rate metric exists and is distinct from the completion-rate metric - Connection reuse (keep-alive) is enabled for any downstream connections the endpoint depends on - TLS certificates on every hop are monitored for expiry, not just configured once
Version Date Change
1.0.0 2026-08-05 Initial publication.