Skip to content

Architecture: Async AI Gateway

Listen to this page11:35
Read the transcript

1. Why direct provider calls don’t scale to production

Host: Let’s start with the simplest possible setup: your application just calls the LLM provider directly. What actually goes wrong with that?

Guest: You inherit every one of that provider’s failure modes with zero protection. If the provider gets slow, every single one of your requests gets slow right along with it. If it goes down, you’re down — fully, with no fallback. And if you’ve got multiple callers sharing that upstream quota, one tenant’s traffic spike can quietly starve everyone else sharing the pipe.

Host: So the fix is putting a gateway in front of all that. What does it actually have to guarantee to be worth the added complexity?

Guest: It has to protect itself, protect the providers behind it, and give every tenant a fair, isolated slice of capacity — all without becoming harder to operate than the problem it’s solving. Concretely that means bounded latency under load, so a caller gets a fast honest rejection instead of hanging forever, bounded blast radius so one bad provider or noisy tenant can’t take everyone else down, and it has to scale horizontally and drain safely on redeploy. That’s the bar we’ll be building against for the rest of this.

2. The order of operations is the architecture

Host: You said the order of operations is the architecture, not an implementation detail. Walk me through what that actually looks like inside the code, because I think people assume request handling is just ‘try the thing, catch errors.’

Guest: Look at generate() in the gateway: it calls _admit() first, and inside _admit, rate limiting happens before the semaphore acquire. That ordering matters because a rejected-for-quota request should never occupy a concurrency slot — that’s capacity a legitimate request needed. Then once admitted, the whole retry loop sits inside a single asyncio.timeout wrapping every attempt, not one timeout per attempt, so a caller’s 15-second budget is a budget for the entire operation, retries included, not 45 seconds nobody asked for.

Host: And what happens to that concurrency slot if something inside the loop blows up — a timeout, a cancellation, a raw exception?

Guest: That’s why the capacity release lives in a finally block wrapping the whole timeout context — it fires no matter how you exit, success, exception, or cancellation. Skip that and you get leaked semaphore slots, which is exactly the leaked-resources failure mode you want to design against. It’s a small detail, but it’s the difference between the code doing what the architecture intends and it quietly drifting from it.

3. The hard failure: a provider that’s slow, not down

Host: So a provider throwing errors is almost the easy case — you can see it, route around it. What’s the scenario that actually breaks people?

Guest: It’s the provider that stays up but gets slow. Nothing fails, no exception to catch, but every request routed there starts dragging. That’s why health-aware selection tracks latency, not just error rate — a technically successful response that’s too slow is still a failure from the caller’s perspective, and you want that provider deprioritized before it drags down everything else.

Host: And that’s where circuit breaking and jittered retries come in, right? Because the naive fix — just retry harder — is the wrong instinct.

Guest: Exactly, immediate synchronized retries across many concurrent requests just amplify the outage — that’s why retries use random jitter instead of a fixed backoff, and why past a failure threshold the circuit breaker stops retrying and fails fast. Same logic applies to the Redis dependency behind rate limiting: if it goes down you either fail closed and protect spend, or fail open with a tighter local limit and protect availability — but you pick one explicitly. The actual failure mode to design against is doing neither, hanging on a Redis call with no timeout and taking the whole gateway down with it.

4. Stateless replicas and the state that secretly isn’t

Host: So you scale the gateway horizontally, spin up more replicas to handle load — and that’s exactly when the rate limiter quietly stops working. Walk me through what breaks.

Guest: The in-process token bucket is per-replica state, but it thinks it’s the whole truth. If you run multiple replicas behind a load balancer, each replica independently enforces the same limit, so the tenant’s real effective quota becomes replica_count times their real limit. Nobody set that limit — it just fell out of how many pods happened to be running that day.

Host: Which means your quota isn’t a business decision anymore, it’s an infrastructure accident. So the fix is moving that state out to Redis — but you were just telling me Redis going down is a designed failure mode in itself.

Guest: Exactly, and that’s the trade you’re actually making, not avoiding. The lab’s Redis limiter does refill-and-consume atomically in one Lua script so concurrent replicas can’t race each other into oversubscribing the bucket — that solves the correctness problem cleanly. But now enforcement has an external dependency, and you’re right back to the fail-open-versus-fail-closed decision, deliberately, deciding what enforcement does when Redis itself becomes unavailable instead of pretending the new dependency can’t fail.

5. Graceful draining and the demo-versus-real identity split

Host: So walk me through the other half of production reliability that nobody thinks about until a deploy goes sideways: what actually happens when SIGTERM lands mid-request?

Guest: The lab’s draining.py has each request register itself as active on entry and deregister in a finally block, so completion is tracked no matter how the request ends. Drain flips a readiness flag to false so the load balancer stops sending new traffic, then waits, bounded by a timeout, for the active count to hit zero. If that timeout fires first, drain returns false explicitly — you get a checkable record that some requests were forcibly cut off, instead of a silent process kill with no idea what was lost.

Host: That’s a clean way to make shutdown honest instead of hopeful. Now, separately — I noticed the lab ships two apps, one with just a header for tenant identity and one with real JWT verification. Isn’t that a security hole?

Guest: It would be if anyone deployed it that way, which is exactly why the docs are blunt about it — the unauthenticated x-tenant-id path exists so you can run bounded concurrency, deadlines, retries, and health-aware routing without also standing up an identity provider just to see the demo work. The secure_app is the one to actually read as the reference: JWT-verified identity, tier-scoped policy at the routing layer, Redis-backed rate limiting shared across replicas. A real deployment collapses these into a single app with the security layer always on — the split is a teaching artifact, not a topology anyone should ship.

6. Streaming changes the cost and observability equation

Host: So we’ve talked about the security split and the demo-versus-real deployment question. Let’s talk about streaming, because I think people assume streaming is just a UX nicety — smoother typing effect — and not something that touches cost or observability at all.

Guest: That assumption falls apart fast at scale. If a user closes the tab three seconds into a ten-second response and the gateway doesn’t notice, it keeps pulling tokens from the provider for the full ten seconds — and every one of those tokens is metered, billed compute for a response literally nobody will see. The fix is checking is_disconnected on every chunk, not just once at the start, because a client can vanish at any point in a long stream. And that’s not a rounding error either — upstream provider cost is almost always the dominant line item in the whole system, way past whatever Redis or the gateway’s own compute costs you, so a lever that stops wasted provider calls has outsized leverage compared to tuning the gateway itself.

Host: And presumably you’re not just fixing the cost, you’re watching for it too — so how does that show up in what you measure?

Guest: Right, we track disconnected as its own counter, completely separate from completed, because a rising disconnect rate is a real signal — are responses too slow and people are giving up, or is some client bug closing connections early? A total-requests-served metric would hide that entirely. And alongside it, time-to-first-token gets its own percentile distribution, at p50, p95, and p99, separate from total stream duration — because ‘streams start slowly’ and ‘streams run long’ are different problems with different fixes, and collapsing them into one duration number just tells you something’s wrong without telling you what.

7. What a real deployment checklist and a fake CI job teach about production readiness

Host: So we’ve talked through the architecture piece by piece — what does it actually look like on the day you ship this thing? What’s on the checklist before someone hits deploy?

Guest: It’s rolling or canary rollout with automatic rollback tied to error-rate and latency SLO breach, and graceful draining wired into the deploy process itself, not just sitting in code somewhere unused. Add pod disruption budgets, topology spread if you’re on Kubernetes, secrets pulled from a real secrets manager and rotated on a schedule, and runbooks for provider outage, Redis outage, a tenant quota incident, and general overload. And on-call ownership written down, not assumed — that last one sounds obvious until the outage happens and three people think someone else owns it.

Host: That’s a good place to end, but I know you’ve got a story about a checklist that looked satisfied and wasn’t — the CI job.

Guest: Right, this is the perfect closer because it’s the whole episode in miniature. There was a Redis integration job that started a real Redis container, then ran the test suite filtered down to just the redis-marked tests — except the only tests matching that filter were backed by a FakeRedis stub that never opens a socket. So the container sat there untouched, and the job was green whether Redis existed or not, including if it had never started at all. It looked exactly like verification and was actually just decoration. The fix does two things: a real test where two limiter instances hit one Redis with double-capacity concurrent acquires, so swapping the atomic Lua script for a naive get-then-set lets forty through instead of twenty — the test measures the atomicity claim instead of restating it. And it sets a flag so a missing Redis fails the build instead of silently skipping, because skipped tests don’t turn anything red, and that’s exactly how you get back to vacuously green a second time, just more quietly. That’s really the whole thesis of this gateway — every green checkmark, every healthy pod, every completed request is a claim, and the job of the architecture is to make sure those claims are actually true, not just comfortable to believe.

Not covered

The planner wanted these and found nothing in the source to support them:

  • Multi-region deployment specifics and per-region quota consistency (raised only as an open interview question, not an implemented architecture)
  • Distributed systems patterns like idempotency keys, vector clocks, and split-brain fencing (Module 2 material, not part of this architecture’s own excerpts)
  • HTTP/2 vs HTTP/3 head-of-line blocking mechanics (Module 3 material not tied to the gateway architecture excerpts directly)

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

An application that calls an LLM provider directly inherits every one of that provider’s failure modes with none of the protection a gateway would add: a slow provider makes every caller slow, a provider outage is a full outage, and nothing stops one caller’s traffic spike from starving every other caller sharing the same upstream quota. The problem this architecture solves: sit a gateway between application traffic and multiple LLM providers that protects itself, protects the providers behind it, and gives every tenant a fair, isolated share of capacity — without becoming so complex that operating it is harder than the problem it solves.

Functional:

  • Route a request to a specific provider (explicit selection) or let the gateway choose the healthiest available one (automatic selection).
  • Support both request/response completions and token-by-token streaming.
  • Resolve tenant identity from verified credentials, not a trusted header alone.

Non-functional:

  • Bounded p99 latency under load — a caller should get a fast, honest rejection before an unboundedly slow success.
  • Bounded blast radius — one misbehaving provider or one noisy tenant must not degrade service for everyone else.
  • Horizontally scalable across stateless replicas, with no single replica as a bottleneck.
  • Safe to redeploy at any time — in-flight requests drain cleanly instead of being killed mid-response.
  • Runs as multiple stateless replicas behind a load balancer; nothing can be safely kept in one replica’s memory if it needs to be correct across all of them (see Module 2’s treatment of this exact constraint for the rate limiter, below).
  • Every LLM provider is a third-party HTTP API with its own rate limits, no uptime SLA the gateway controls, and latency the gateway cannot influence — only route around.
  • No single piece of supporting infrastructure (a Redis instance, say) may become a hard dependency whose outage is a full gateway outage; every such dependency needs an explicit degraded-mode behavior.
  • Tenant identity in the production-hardened path comes from a verified JWT — an easily-spoofed header is acceptable only for the unauthenticated demo path, never the real one.

The order of these steps is the architecture, not an implementation detail: identity is verified before quota is checked (an unauthenticated caller shouldn’t consume a tenant’s quota), quota is checked before admission control (a rejected-for-quota request shouldn’t occupy a concurrency slot), and the deadline wraps the entire retry loop rather than each individual attempt (a caller’s timeout budget is a budget for the whole operation, retries included) — see Module 1’s Architecture section for the code-level version of this same reasoning.

A provider degrades without fully failing

The harder case is not a provider returning errors — that’s easy to detect — but a provider that stays up while getting slower. Health-aware selection tracks latency, not just error rate, specifically so a “technically successful but too slow” provider gets deprioritized before it drags down every request routed to it.

Redis becomes unavailable

If distributed rate limiting depends on Redis and Redis goes down, the gateway must make an explicit choice, not an accidental one: fail closed (protect spend and quota controls, at the cost of availability) or fail open with a tighter local emergency limit (protect availability, at the cost of weaker quota enforcement). Silently doing neither — crashing, or hanging on a Redis call with no timeout — is the actual failure mode to design against.

One tenant's traffic spike starves another tenant

Without per-tenant isolation, a single tenant’s burst can consume the shared concurrency pool and shared provider quota that every other tenant also depends on. Per-tenant rate limiting (see Trade-offs, below) exists specifically to contain this.

A rolling deploy kills in-flight requests

A replica that receives SIGTERM and exits immediately drops every request it was mid-processing. Graceful draining — stop accepting new traffic, wait for active requests to finish within a bounded timeout, then exit — is what turns a routine deploy from a guaranteed blip into a non-event. See Module 1’s Production Example for the actual draining implementation this depends on.

  • Stateless replicas scale horizontally by design — the harder scaling question is what state isn’t actually stateless. An in-process, per-replica rate limiter is the concrete example: it scales the wrong way, silently multiplying each tenant’s effective quota by the replica count (replica_count × their real limit) — the exact trade-off Module 2 walks through in detail.
  • Outbound connections to providers can hit ephemeral-port or connection-pool ceilings at high fan-out, particularly behind NAT — a scaling failure mode that shows up as intermittent connection errors under load, not a clean capacity error (see Module 3’s Scaling section).
  • Autoscale on saturation, not CPU. Queue-wait time and concurrency-slot saturation are leading indicators of overload for an I/O-bound gateway; CPU usage tends to stay flat right up until the system falls over.
  • Tenant identity from a verified JWT (issuer, audience, expiry, key rotation via JWKS in production) — never a client-supplied header alone.
  • Tenant-tier-scoped provider and model access policy, enforced at the routing layer, not left to the caller’s honesty.
  • Secrets (JWT signing keys, Redis credentials, provider API keys) injected at runtime from a secrets manager or environment, never committed to configuration.
  • Redis connections authenticated and encrypted in transit in any real deployment — an internal cache is still a credential and quota store, not a throwaway component.
  • Audit events emitted for policy denials and any privileged provider selection, so “who was denied what, and why” is answerable after the fact.

A demo entry point vs. a security-hardened entry point

Keeping an unauthenticated demo path (x-tenant-id header, no verification) separate from the JWT-and-policy-aware production path keeps the reliability patterns (bounded concurrency, deadlines, retries, health-aware routing) legible on their own, without every example also threading through an identity provider. A real deployment collapses these into one app with the security layer always on — the split exists for teaching clarity, not as a recommended production topology.

In-process vs. distributed rate limiting

Covered in full in Module 2: in-process is simpler and has no external dependency, but is wrong the moment there’s more than one replica. Distributed (Redis-backed) rate limiting is correct across replicas at the cost of a new dependency and a new failure mode — what enforcement does when that dependency is unavailable (see Failure Modes, above).

Automatic provider fallback vs. an explicit provider request

Automatic fallback maximizes availability — if the requested provider is unhealthy, try another. But a caller who explicitly requested a specific provider (for cost, compliance, or data-residency reasons) should never be silently rerouted; this architecture only auto-falls-back when no specific provider was requested, treating an explicit request as an intent that overrides availability optimization.

  • Upstream provider cost (tokens/API calls) is almost always the dominant line item — far exceeding the compute cost of running the gateway itself. Architecture decisions that reduce wasted provider calls have outsized cost leverage compared to gateway-side optimization.
  • Circuit breaking avoids paying for calls to a provider that’s already failing — every call a broken circuit prevents is a call that would have been billed and still failed.
  • Disconnect detection on streaming responses avoids paying for abandoned work — a client that closes its connection mid-stream shouldn’t keep the gateway pulling (and paying for) tokens no one will see; see Module 3’s Production Example for the exact mechanism.
  • Redis and observability infrastructure are real but secondary costs — worth optimizing only after the provider-cost and wasted-work levers above have been addressed.
  • RED metrics (rate, errors, duration) per provider and per tenant, not just globally — a single aggregate latency number hides which tenant or provider is actually degraded.
  • Queue-wait time tracked separately from provider execution time — otherwise a system stuck queueing for concurrency slots looks identical to one that’s genuinely slow to process (see Module 1’s Failure Modes section).
  • Distributed traces spanning identity verification, quota decision, routing, retries, and the outbound provider call — enough to answer “where did this specific slow request actually spend its time” without guessing.
  • Streaming-specific telemetry: time-to-first-token and stream duration as separate percentile distributions, plus a disconnect-rate counter distinct from the completion-rate counter (see Module 3’s Implementation section).
  • Structured logs correlated by request and trace ID, deliberately excluding prompt or response content to avoid leaking sensitive data into log storage.

Before this architecture is production-ready

  • Rolling or canary rollout with automatic rollback criteria tied to error-rate/latency SLO breach
  • Graceful draining wired into the deploy process, not just implemented in code - Pod disruption budgets and topology spread configured if running on Kubernetes - Secrets (JWT keys, Redis credentials, provider API keys) sourced from a secrets manager, rotated on a schedule - Runbooks written for: provider outage, Redis outage, a tenant quota incident, and general overload - On-call ownership and escalation path documented, not assumed

This checklist mirrors the lab’s own PRODUCTION_READINESS.md — treat an unchecked item there as a known gap against this architecture, not an oversight.

Design a multi-tenant gateway in front of several LLM providers, from scratch.

A strong answer walks the Request Flow order above and explains why that order (identity before quota, quota before admission, deadline wrapping retries) rather than just listing the components. Weak answers list technologies (Redis, JWT, circuit breaker) without explaining the ordering that makes them work together correctly.

How do you decide fail-open vs. fail-closed when a supporting dependency like Redis goes down?

Anchor the answer to what the dependency protects: if it’s protecting spend or hard abuse limits, fail closed. If it’s protecting a soft quota where availability matters more than perfect enforcement, fail open with a tighter local emergency limit. The wrong answer is “it depends” with nothing after it — name the concrete factor that decides it.

How would this architecture change to serve multiple regions?

Expect a candidate to raise: per-region gateway deployments to keep provider calls low-latency, what tenant quota state needs to be globally consistent versus region-local, and how health-aware provider selection interacts with region-specific provider availability and pricing. There’s no single right answer — the interview signal is whether the candidate identifies which pieces of this architecture are the ones that get harder, rather than assuming multi-region is a free scaling dimension.

Hands-on Lab

This architecture is backed by a real, tested implementation, not a diagram alone. Read the lab documentation →

labs/async-ai-gatewayproduction-ready