Async AI Gateway
Read the transcript
1. Two apps, one core: what the gateway actually does
Host: Welcome back. Today we’re cracking open a lab called the Async AI Gateway, and this one’s a great teaching tool because it doesn’t just show you the happy path — it shows you where a supposedly solid test turns out to be hollow, and where real engineering judgment has to step in. So let’s start with the basics: what are we actually looking at when we open this repo?
Guest: So there are two FastAPI apps sitting on top of one shared core called AsyncAIGateway. The first, production_app, gives you bounded concurrency, deadlines, retries with jitter, health-aware fallback between providers, RED metrics, and OpenTelemetry hooks — it identifies tenants with a simple header so you can run it without standing up an identity provider. Then there’s secure_app, which is everything production_app has, plus JWT-verified tenant identity, tier-based provider policy, and Redis-backed rate limiting that’s shared correctly across replicas. That second one is really the reference for what a real deployment should look like.
Host: Why split them at all instead of just shipping one app with security always on?
Guest: Purely pedagogical — it keeps the identity and quota layer legible on its own instead of threading JWT checks through every single reliability example. In production you’d collapse them into one app with security always on. And it’s genuinely easy to poke at locally: you set up the virtualenv, pip install the extras, run uvicorn on production_app, and you can curl the generate endpoint with a tenant header, or just run docker compose up and get the gateway, Redis, and an OpenTelemetry collector all wired together at once.
2. The Redis test that was green for the wrong reason
Host: Okay, so let’s talk about the part of this lab that I think is the most humbling story: the Redis integration job that was green for months and testing basically nothing. What was actually going on there?
Guest: So the CI job would spin up a real Redis service container, then run pytest with a filter, dash-k redis, to select the relevant tests. The problem is the only tests matching that filter were backed by a FakeRedis stub — an in-memory fake that returns canned answers and never opens a socket. The container just sat there unused, and the job was green whether Redis existed or not.
Host: So it would have passed even if the container had never started at all. How do you even catch that kind of thing, and what did the fix look like?
Guest: Right, that’s the scary part — it’s a passing test with zero signal. The fix was a new test file that actually talks to a live server: two limiter instances sharing one Redis, and you fire two times capacity acquires concurrently at a single tenant. Exactly capacity should succeed, and if you swap the atomic Lua script for a naive HGET-then-HSET, you watch 40 of 40 get through where the atomic version correctly allows only 20 — so the test measures the atomicity claim instead of just restating it. We also set REDIS_INTEGRATION_REQUIRED equals 1, so a missing REDIS_URL fails the build instead of quietly skipping, because skipped tests don’t turn anything red and the job could go vacuous again just as silently.
3. Where the demo still falls short of production
Host: Okay, so the test is fixed and honest now. But zoom out — what does this lab actually teach about the gap between a working demo and something you’d trust in production?
Guest: A bunch of things that look like nitpicks but aren’t. Semaphores cap concurrent work, token buckets cap arrival rate, and you actually need both enforced at different points, not one standing in for the other. Then there’s the atomicity point we just proved — naive GET-then-SET against Redis lets 40 through at a capacity of 20, only the Lua script gets it right. And health-based routing has a nasty feedback loop hiding in it: score a provider unhealthy, starve it of probe traffic, and it can never prove it recovered, so you need decay and caps on those adjustments. Streaming makes this worse because once a partial response has hit the client you can’t silently retry — the caller’s already rendering it. And readiness versus liveness is the one that bites people in rollouts: readiness says can this pod safely take traffic, liveness says should this process be restarted, and conflating them means your rolling deploy kills pods that were just draining correctly.
Host: So none of that is fixed in this lab yet — it’s sitting in the readiness doc as known work, not something someone forgot.
Guest: Exactly, that’s the point of tracking it in the production readiness document instead of pretending it’s done — unchecked items are deliberate gaps across correctness, reliability, observability, tenancy, capacity, delivery. That’s the principal-level habit this whole lab is really modeling: distinguish what you’ve actually verified, like that Redis test now does, from what you’ve simply not broken yet. Demo versus production is exactly that line.
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
An async API gateway that sits in front of one or more LLM providers and protects both itself and those providers under load. It runs with deterministic fake providers out of the box — no API keys required — while the adapters underneath stay vendor-neutral.
Source: labs/async-ai-gateway
Entry points
Section titled “Entry points”The lab ships two FastAPI apps that layer capability on top of a shared AsyncAIGateway core:
ai_gateway.production_app— bounded concurrency, deadlines, retries with jitter, health-aware provider fallback, RED metrics, and OpenTelemetry hooks. Tenant identity comes from anx-tenant-idheader, which keeps the demo runnable without an identity provider.ai_gateway.secure_app— everything inproduction_app, plus JWT-verified tenant identity, tier-based provider policy, and Redis-backed distributed rate limiting shared across replicas. This is the entry point to read as the reference for a real deployment.
Request flow
Section titled “Request flow”flowchart TD
Client[Client] --> Auth[JWT identity verification]
Auth -->|invalid| Reject401[401 Unauthorized]
Auth -->|valid| Quota[Tenant quota check\nRedis token bucket]
Quota -->|exceeded| Reject429[429 Too Many Requests]
Quota -->|allowed| Admission[Bounded concurrency\nadmission control]
Admission -->|saturated| Reject503[503 Service Unavailable]
Admission -->|admitted| Telemetry[Request telemetry\nID, RED metrics, OTEL span]
Telemetry --> Gateway[AsyncAIGateway\ndeadline + retry budget]
Gateway --> Health[Health-aware provider selection]
Health --> Provider1[Primary provider]
Health -->|ejected on failure| Provider2[Fallback provider]
Provider1 -->|transient failure, retries remain| Gateway
Provider2 --> Response[Response or SSE stream]
Provider1 --> Response
Response --> ClientWhat it demonstrates
Section titled “What it demonstrates”- Bounded concurrency and queue-wait limits, separate from rate limiting.
- Request deadlines, cancellation, retries, and jitter within a fixed retry budget.
- Circuit breaking and ordered provider fallback, with explicit-provider requests never silently falling back.
- A Redis-backed token bucket that executes refill-and-consume atomically in one Lua script, so concurrent replicas can’t oversubscribe a tenant’s quota.
- JWT-verified tenant identity (
sub,tenant_id,tierclaims) with tier-scoped provider and timeout policy. - Health-aware provider selection based on success rate, consecutive failures, and a latency EMA, with automatic ejection and recovery.
- Graceful connection draining on shutdown, and SSE streaming that stops work on client disconnect.
- Correlated structured logs, RED metrics, Prometheus output, and OpenTelemetry trace export.
Run it
Section titled “Run it”cd labs/async-ai-gatewaypython3.12 -m venv .venvsource .venv/bin/activatepip install -e '.[dev,observability,redis]'uvicorn ai_gateway.production_app:app --reloadcurl -s http://127.0.0.1:8000/v1/generate \ -H 'content-type: application/json' \ -H 'x-tenant-id: tenant-demo' \ -d '{"prompt":"Explain bounded concurrency"}'The full local stack — gateway, Redis, and an OpenTelemetry Collector — comes up with:
docker compose up --buildVerify it
Section titled “Verify it”pytestruff check .mypy srcGitHub Actions runs this same check, plus a Redis-integration job, on every change under
labs/async-ai-gateway/ — see
.github/workflows/async-ai-gateway-ci.yml.
The integration job that was testing nothing
Section titled “The integration job that was testing nothing”That Redis job started a service container and then selected tests with pytest -k redis. The
only tests matching were the ones backed by a FakeRedis stub, which return a canned answer and
never open a socket. The container was never touched, and the job was green whether or not Redis
existed — including if it had never started.
The replacement, tests/test_redis_integration.py, talks to a real server, and its central test is
the one a fake cannot express: two limiter instances sharing one Redis, 2 × capacity acquires
fired concurrently at a single tenant. Exactly capacity may succeed. Swapping the Lua script for
a naive HGET-then-HSET lets 40 of 40 through where the atomic version allows 20 — so the
test measures the atomicity claim rather than restating it.
The subtler half of the fix is that the job now sets REDIS_INTEGRATION_REQUIRED=1, and under that
flag a missing or unreachable REDIS_URL fails rather than skips. Without it the job could
regress to vacuously green a second time, just more quietly: skipped tests do not turn a build red.
Keying that off CI would not work — GitHub Actions sets CI for every job, including the one
with no Redis attached.
Production readiness
Section titled “Production readiness”The lab tracks its own gap list against a real production bar in
PRODUCTION_READINESS.md —
covering correctness contracts, reliability, observability, tenancy and security, capacity, and
delivery. Treat an unchecked item as a known gap, not an oversight.
Principal-level discussion points
Section titled “Principal-level discussion points”- Semaphores bound concurrent work; token buckets bound arrival rate. Production gateways generally need both, enforced at different points in the request path.
- A distributed quota decision needs atomic state mutation — a naive
GETthenSETagainst Redis oversubscribes under concurrency, which is why the limiter uses a single Lua script. Measured at a capacity of 20 with two replicas: the script allows 20, the naive version allows all 40. - Health-based routing can create feedback loops: a provider that looks unhealthy under one scoring window can starve of the probe traffic it needs to recover. Production systems cap scoring adjustments and expire stale observations.
- Streaming fallback is unsafe once partial output has reached the client — you cannot silently retry a stream the caller has already started rendering.
- Readiness reflects whether traffic can be accepted safely; liveness only reflects whether the process should be restarted. Conflating them causes rolling deploys to kill draining pods.