Skip to content

Module 4: AI Infrastructure

Listen to this page14:31
Read the transcript

1. What AI infrastructure actually is

Host: Welcome to Module 4. We’ve spent the last two modules on distributed systems and networking fundamentals — the stuff that underlies basically any backend. Today we’re narrowing in on something more specific: what happens when the thing you’re calling isn’t a database or a microservice, but a model.

Guest: Right, and that’s really the whole premise of AI infrastructure — it’s not machine learning research, and it’s not generic backend plumbing either. It’s the layer in between that takes a model, whether it’s someone else’s API or one you’ve trained yourself, and turns it into a product feature that’s actually reliable, affordably priced, and safe to change over time. Think of this module as the bridge between the general systems material we’ve already covered and the AI-specific protocols and frameworks we’ll get into later — it’s about the concerns that show up precisely because your dependency is a model, not a server you fully control.

2. Layers that change at different speeds

Host: So if the models and providers are the fast-moving piece, how does that actually change the way you structure the stack? Like, where does that speed mismatch show up architecturally?

Guest: Think of three layers moving at different clock speeds. Models and providers churn constantly — new versions, new vendors, pricing shifts, sometimes the same API quietly gets worse or better underneath you. The gateway layer changes occasionally, maybe you add a policy or a new provider. And the application, the actual product surface and prompts, changes slowest of all. The design principle is simple: model and provider choice has to be a routing decision made at runtime, not something hardcoded into the slow-changing layer.

Host: So if swapping a provider means editing application code, that’s a signal the dependency arrow is pointing the wrong way.

Guest: Exactly, and that’s precisely why in the async-ai-gateway lab, provider selection happens in a select function, evaluated per request, instead of being a compile-time import somewhere in the app. Teams that wire an SDK straight into product code are betting they’ll never need a second provider or tenant-based routing, and that bet almost always fails within a year. The gateway isn’t extra ceremony — it’s the thing that absorbs the fast-changing layer before it spreads everywhere.

3. Why ‘200 OK’ isn’t enough

Host: So what happens when the system answers just fine, and the answer is just… wrong? A 200 status code doesn’t tell you that anything went bad.

4. Build vs. buy, and non-determinism as an engineering fact

Guest: Right, that’s the evaluation gap, and it connects to something deeper: the same prompt can produce different outputs on different calls. That’s not a bug, it’s the nature of these models, but it breaks the core assumption every other testing discipline relies on, that a system under test behaves deterministically.

Host: So the unit test playbook just doesn’t apply here. You can’t assert that the output equals some fixed string.

Guest: Exactly, exact-match assertions are the wrong tool entirely. What you need instead is an evaluation harness, a fixed suite of representative prompts scored against a rubric or reference set, run every time you touch a prompt or swap a model, the same way you’d run a regression suite before shipping code.

Host: Which reframes something people tend to treat casually, editing a prompt in production. That’s not a config tweak, that’s a deploy, and it sounds like it needs the same guardrails, canary a slice of traffic, compare against baseline, have a rollback path.

5. Cost per request, not per month

Host: Let’s talk money, because I think most teams still think of AI spend the way they think of a cloud bill — check it at the end of the month, wince, move on. But you’re saying that’s already too late to catch the actual problem.

Guest: Right, because the pricing structure itself is asymmetric in a way that a monthly total completely hides. Providers charge separately for input and output tokens, and output tokens are typically several times more expensive than input. So ‘summarize this document’ is cheap — big input, tiny output — but ‘expand this outline into a full report’ is expensive, small input, large output. Two features can look similar in your product and have wildly different cost profiles per call, and if you’re only looking at a monthly total, one expensive feature can quietly dominate your entire spend with zero visibility into which one it was.

Host: So you need cost attached to the individual request, not the invoice. What does that actually look like in code — is this just logging a number somewhere?

Guest: It’s a small wrapper, but the details matter. You define a usage record — tenant, provider, input tokens, output tokens, latency — and estimated cost is a property computed on demand from a pricing table, not baked in at write time. That means two things: if a provider corrects their pricing retroactively, you can re-price historical records instead of having stale numbers frozen forever, and the pricing table stays the single source of truth instead of drifting across every call site that happens to hardcode a rate. Then a function that wraps the actual call times it and emits that record — so every single request produces a cost and latency data point you can attribute to a tenant.

6. A model swap is a deploy

Host: So you’ve got requests flowing through this wrapper, cost and latency attributed per tenant. But at some point someone’s going to say ‘let’s just switch our default provider, Provider B is cheaper.’ Is that a bigger deal than it sounds?

Guest: It’s exactly as big a deal as a code deploy, because that’s what it is. The naive version is you change one line in a config, deploy, and now a hundred percent of traffic is hitting a model that might phrase things worse, or just behave differently on your specific prompts — and nobody notices until support tickets pile up, because there’s no error, just quieter degradation.

Host: So what does doing it right actually look like, end to end?

Guest: Same shape as any canary rollout. Route five percent of traffic to Provider B, run your evaluation harness’s rubric against both arms for a fixed window, and compare quality score, p95 latency, and cost per request side by side. If B holds up you ramp to twenty-five, then a hundred; if quality drops below your threshold at any stage, you roll back to A. The point is the decision is made from data you already have sitting there, not from someone’s gut feeling after skimming a few outputs.

7. How this quietly breaks: failure modes and security

Host: So say the rollout looks clean, you’ve ramped to a hundred percent, everyone moves on. Is that actually the end of the story, or does this stuff keep breaking quietly after the fact?

Guest: It keeps breaking, that’s the uncomfortable part. A provider can update the model behind that same stable endpoint months later, and quality shifts with zero error code, nothing a normal monitor flags — only a running eval harness comparing against baseline catches it. Same with cost: a request with no cap on output length, or a prompt that nudges the model into a long ramble, quietly turns into an expensive request, and since output tokens cost more than input, that adds up fast if you’re not capping length explicitly.

Host: And I’d guess your eval suite itself can lull you into false confidence if it’s only testing the easy cases.

Guest: Exactly — happy-path-only evals miss adversarial inputs, weird formatting requests, the actual tail of real traffic, which is where production failures live. That same untrusted-input problem is also a security seam: a model summarizing a document or reading a webpage can follow instructions buried in that content, which shows up first as ‘broken’ behavior in your monitoring before anyone calls it prompt injection — so you separate trusted instructions from untrusted data, scope what context and tools each call can reach, and treat prompts and logs with the same PII discipline as any other system touching personal data. On top of that, watch provider rate limits as their own capacity constraint — your gateway’s admission control does nothing once the ceiling lives on their side.

8. The cost-latency-quality triangle and scaling routing problems

Host: So let’s talk performance trade-offs, because I don’t think ‘make it faster’ is actually one problem. Time-to-first-token and tokens-per-second sound related but you’re describing them as almost separate levers.

Guest: They are separate. Time-to-first-token is about perceived responsiveness — did the thing start talking to me — which matters enormously for a chat interface, per Module 3’s discussion of latency. Tokens-per-second throughput is about total duration for long outputs, which matters more for something like a report generation job nobody’s staring at live. And both sit inside this cost-latency-quality triangle: cheaper, faster models are usually lower quality, so pushing on speed or price pushes back on the third corner. That has to be a measured, per-use-case decision, not one global default model for everything — and caching, semantic or exact-match, is a real lever on both cost and latency that we get into properly in the semantic cache lab.

Host: Okay, and then that per-use-case routing decision has to survive actual scale — multiple providers, real traffic. What breaks first?

Guest: Quota management, mostly — at scale you need to know each provider’s remaining quota and route proactively, not just fall back reactively after a 429. And fallback chains need to scale with traffic, not sit there as a static ordered list — a fallback sized for occasional overflow can get overwhelmed itself if the primary degrades under a broad outage, which is exactly when everyone’s traffic shows up at its door. Then multi-region adds another dimension, because not every model is available or priced the same everywhere, so scaling to a new region turns into a routing-policy problem, not just spinning up infrastructure.

9. The gateway lab as the routing layer made real

Host: So everything we’ve just described in the abstract — health-aware selection, quota-aware routing, fallback that scales — is there an actual reference implementation we can point to, rather than just a mental model?

Guest: Yes, that’s exactly what the async-ai-gateway lab is for. It ships two apps on top of a shared core: the production app has bounded concurrency, deadlines, retries with jitter, and health-aware fallback, while the secure app adds JWT-verified tenant identity, tier-based provider policy, and a Redis-backed rate limiter shared across replicas. The split exists so the identity and quota layer is legible on its own, though a real deployment would collapse them into one app with security always on.

Host: That Redis limiter is interesting given what you said about quota tracking at scale — what actually makes it safe under concurrency, where a naive counter wouldn’t be?

Guest: It’s a single Lua script doing refill-and-consume atomically, so two replicas can’t both read the same remaining quota and each think they’re clear to spend it — in testing, a capacity of 20 with two replicas let through exactly 20 with the script versus all 40 with a naive get-then-set. And it’s not the only subtlety: health scoring can create feedback loops where an unhealthy-looking provider starves of the probe traffic it needs to recover, streaming fallback is unsafe once a client has started rendering partial output, and conflating readiness with liveness is what makes rolling deploys kill pods that were just draining connections, not actually broken.

10. When the test suite lies: the Redis integration story

Host: You mentioned testing earlier, but I want to end on this because it’s such a good gut-punch story. There was a Redis integration job in CI that was green the entire time — tell me how that happened.

Guest: The job spun up a Redis service container and then ran the test suite filtered down to anything with ‘redis’ in the name, but the only tests matching that filter were backed by a FakeRedis stub — canned answers, no socket ever opened. So the container just sat there unused, and the job was green whether Redis existed or not, including if it had never started at all. The fix, a new Redis integration test file, does the one thing a fake can’t: two limiter instances against one real Redis, twice the capacity fired concurrently, and it asserts exactly capacity gets through — swap the Lua script for a naive get-then-set and you watch it let 40 of 40 through instead of 20. The subtler half is a required-integration flag that turns a missing Redis connection URL into a hard failure instead of a skip, because a skip doesn’t turn the build red — it just goes vacuously green again, more quietly than before.

Host: So the lesson isn’t ‘write more tests,’ it’s ‘make sure the test can actually fail for the reason you think it can.’ That feels like a fitting note to close the whole module on — treat gaps as known and tracked, not hidden behind something that looks like coverage.

Guest: Exactly, and that’s why the lab keeps its own production readiness list right next to the code — correctness contracts, reliability, observability, tenancy, capacity, delivery — with unchecked items left visibly unchecked. An honest gap list beats a green checkmark that isn’t measuring anything, and that’s the whole discipline of this layer in one sentence: assume non-determinism, verify your guarantees under real conditions, and never let ‘it passed CI’ stand in for ‘it works in production.’

Not covered

The planner wanted these and found nothing in the source to support them:

  • Detailed walkthrough of the planned Model Router and Semantic Cache labs’ internal design (only mentioned as future work, not detailed)
  • Comparison to specific competing gateway products or vendors beyond the generic Provider A/B example
  • A deep dive into the cited external references’ full arguments (Huyen, Willison, Yan) beyond their one-line framing

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

“AI infrastructure” isn’t machine learning research, and it isn’t generic backend engineering — it’s the stack in between: the layer that takes a model (someone else’s API or your own) and turns it into a product feature that’s reliable, affordably priced, and safe to change over time. This module is the bridge between the general distributed-systems and networking material in Modules 2–3 and the AI-specific protocols and frameworks in Modules 5–7: it covers the concerns that show up specifically because the “backend” you’re calling is a model, not a database.

An AI platform stack has layers that change at very different rates, and the architecture should reflect that mismatch, not fight it:

  • Models and providers change fastest — new versions, new vendors, pricing changes, sometimes silent quality shifts underneath a stable API.
  • The gateway/routing layer (Modules 1–3’s material) changes occasionally — new policies, new providers added.
  • The application changes slowest — the product surface and prompts built on top.

The core design principle: treat model and provider choice as a routing decision, not a hardcoded dependency. If swapping the fastest-changing layer requires touching the slowest-changing one, the architecture has the dependency direction backwards. This is exactly why labs/async-ai-gateway’s provider selection is a runtime decision (_select(), evaluated per-request) rather than a compile-time import — the gateway exists specifically to absorb the layer that changes fastest.

The gateway is the seam, not an afterthought

Teams that bolt a provider SDK directly into application code are making a bet that they’ll never need to swap providers, add a second one, or route by tenant/cost/latency. That bet is almost always wrong within a year. The gateway layer this module assumes exists isn’t extra infrastructure — it’s where the “models change fast” reality gets contained before it spreads.

The evaluation and observability layer wrapping every other layer is deliberate, not decorative: unlike a typical backend where “did it work” is a status code, an AI system’s failure mode is often “it returned 200 with a wrong or low-quality answer” — invisible to conventional monitoring unless something is specifically watching for it.

Build vs. buy. Hosted provider APIs win on time-to-market, no infrastructure to operate, and access to frontier model quality without a training budget. Self-hosting wins on cost at sufficient scale, latency control (no third-party network hop), data residency and compliance requirements, and the ability to fine-tune. The crossover point is a real number, not a vibe — model it explicitly: hosted API cost scales linearly with usage; self-hosting has a large fixed cost (GPUs, ops headcount) and a much lower marginal cost per request. Below the crossover volume, hosted wins; above it, self-hosting wins, assuming the operational capability to run it exists.

Non-determinism is a first-class engineering concern, not a research detail. The same prompt can produce different outputs on different calls — this breaks the assumption every other module in this handbook relies on, that a system under test behaves deterministically. Unit tests that assert exact output equality are the wrong tool; the right tool is an evaluation harness that scores outputs against a rubric or reference set, run against a fixed test suite on every prompt or model change — the AI-specific equivalent of a regression test suite.

Cost modeling. Hosted providers typically price input and output tokens separately, with output tokens usually costing several times more than input tokens — a detail that changes the economics of, say, “summarize this document” (cheap: large input, small output) versus “expand this outline into a report” (expensive: small input, large output). Cost needs to be tracked per request, not just as a monthly bill total, or a single expensive feature can silently dominate spend with no way to attribute it.

A model or prompt change is a deploy, and deserves the same rigor as one. Swapping a default model, bumping a provider’s model version, or editing a production prompt changes user-facing behavior exactly the way a code deploy does — and deserves the same tools: canary a percentage of traffic, compare quality/cost/latency against the baseline, and have a rollback path. Treating a prompt edit as “just a config change” with no evaluation gate is how quality regressions ship silently.

Research Note

One of the clearest treatments of why traditional software engineering practices (testing, monitoring, deployment) need real adaptation — not abandonment — for AI-backed systems. Directly informs this module’s Deep Dive and Failure Modes sections.

Source: Chip Huyen, "Building LLM Applications for Production"

A minimal cost-tracking wrapper — the concrete mechanism behind “track cost per request, not just per month” from the Deep Dive section:

Per-request cost tracking

from __future__ import annotations
import time
from collections.abc import Awaitable, Callable
from dataclasses import dataclass
# Illustrative pricing table. Real systems load this from a versioned config that
# changes independently of any code deploy — provider prices change on their schedule,
# not yours.
TOKEN_PRICE_PER_1K = {
"provider-a/large": {"input": 0.003, "output": 0.015},
"provider-b/small": {"input": 0.0005, "output": 0.0015},
}
@dataclass
class UsageRecord:
tenant_id: str
provider: str
input_tokens: int
output_tokens: int
latency_ms: float
@property
def estimated_cost_usd(self) -> float:
price = TOKEN_PRICE_PER_1K.get(self.provider)
if price is None:
return 0.0
return (
(self.input_tokens / 1000) * price["input"]
+ (self.output_tokens / 1000) * price["output"]
)
async def track_usage(
tenant_id: str,
provider: str,
call: Callable[[], Awaitable[tuple[str, int, int]]],
emit: Callable[[UsageRecord], None],
) -> str:
started = time.perf_counter()
text, input_tokens, output_tokens = await call()
emit(UsageRecord(
tenant_id=tenant_id, provider=provider,
input_tokens=input_tokens, output_tokens=output_tokens,
latency_ms=(time.perf_counter() - started) * 1000,
))
return text

estimated_cost_usd is a property, computed on demand from the pricing table, rather than baked in at call time — so historical usage records can be re-priced if the pricing table is corrected, and so the pricing table itself stays the single source of truth instead of drifting across call sites.

A team wants to switch their default provider from Provider A to a cheaper Provider B. Done naively (“just change the config and deploy”), a quality regression reaches 100% of traffic before anyone notices. Done per this module’s Deep Dive: route 5% of traffic to Provider B, run the evaluation harness’s rubric against both arms’ outputs for a fixed period, compare quality score, p95 latency, and cost per request side by side. If Provider B holds up, ramp to 25%, then 100%; if quality drops below a defined threshold, roll back to 100% Provider A — the same canary-and-rollback shape any other production deploy would use, applied to a model swap instead of a code change.

Silent model degradation

A provider updates the model version behind a stable API endpoint, and output quality shifts — sometimes for the worse — with no error, no status code change, nothing a conventional monitor would catch. Only a running evaluation harness comparing current output quality against a baseline can detect this.

Cost runaway from unbounded generation length

A request with no max_tokens cap, or a prompt that induces unusually long output, can turn a routine request into a disproportionately expensive one — especially given output tokens’ higher price per the Deep Dive section. Cap output length explicitly; don’t rely on the model to stop at a “reasonable” length.

Evaluation blind spots

An evaluation suite that only tests happy-path prompts misses the failure modes that actually matter in production — adversarial inputs, edge-case formatting requests, and the tail of real user queries the design didn’t anticipate. An evaluation suite is only as good as its coverage of what real traffic actually looks like.

Prompt injection is a reliability issue, not just a security one

A model that follows instructions embedded in untrusted input (a document it’s summarizing, a webpage it’s reading) can be steered into unexpected behavior even without malicious intent behind the input — this shows up as “broken” behavior in evaluation and production monitoring well before anyone frames it as a security finding. See Security, below, for the mitigation side.

Vendor rate-limit surprises during traffic spikes

A provider’s rate limit that was never a constraint at normal traffic becomes a hard ceiling during a spike, with the gateway’s own admission control (Module 1) doing nothing to prevent it — the constraint lives on the provider’s side, not the gateway’s, and needs to be modeled and monitored as its own capacity limit.

Hosted API vs. self-hosted model

Hosted wins on speed to ship and zero infrastructure ownership; self-hosted wins on cost at scale, latency control, and compliance/data-residency requirements. Model the crossover volume explicitly (Deep Dive) rather than defaulting to whichever a team is more comfortable with.

Synchronous vs. sampled evaluation

Evaluating every production response in real time is the most thorough option but adds latency and cost to every request. Sampling a percentage of traffic for evaluation is cheaper and faster but trades completeness for cost — appropriate once a system has enough volume that a representative sample is statistically meaningful.

One large general model vs. several smaller specialized ones, routed by task

A single large model is simpler to operate and reason about. Multiple smaller, cheaper models routed by task (classification-style requests to a cheap fast model, complex reasoning to an expensive capable one) can meaningfully cut cost and latency at the expense of routing complexity and more evaluation surface area to maintain.

  • Prompt injection: untrusted content a model processes (documents, web pages, tool outputs) can contain instructions that steer its behavior. Mitigate with clear separation between trusted instructions and untrusted data in the prompt structure, and least-privilege on what tools or data a given model call can actually reach.
  • Data exfiltration via model output: a model with access to sensitive context and an output channel (even just its response text) can be tricked into leaking that context — scope what context any single call has access to, matching the actual task’s need.
  • PII in prompts and logs: prompts and responses often contain user-supplied personal data; logging and observability pipelines need the same PII handling discipline as any other system that touches that data, not an exemption because “it’s just a prompt.”
  • Time-to-first-token vs. tokens/sec throughput are different numbers that matter for different reasons — the former drives perceived responsiveness (Module 3), the latter drives total request duration for long outputs.
  • The cost-latency-quality triangle: cheaper and faster models are usually lower quality, and pushing on any one corner costs on the other two. Make this trade-off an explicit, measured decision per use case rather than a single global default model.
  • Caching — semantic or exact-match caching of repeated or similar requests is a real lever for both cost and latency, covered in depth by the planned Semantic Cache lab (see Hands-on Lab).
  • Provider quota management across multiple providers becomes a real scheduling problem at scale — knowing each provider’s remaining quota and routing accordingly, not just falling back reactively after a 429.
  • Fallback chains need to scale with traffic, not just exist as a static ordered list — a fallback provider sized for occasional overflow can itself become overwhelmed if the primary degrades under the exact conditions (a broad outage) that push traffic to it.
  • Multi-region model availability — not every provider or model is available, or priced the same, in every region, which turns “scale to a new region” into a routing-policy problem, not just an infrastructure one.

How do you decide between a hosted API and self-hosting a model?

Model the crossover point explicitly: hosted cost scales linearly with usage, self-hosting has a large fixed cost and lower marginal cost. Below the crossover volume hosted wins; above it, self-hosting wins if the operational capability exists. Compliance and latency-control requirements can override the pure cost calculation either direction.

How would you detect a provider silently degrading model quality behind a stable API?

A running evaluation harness comparing current output against a quality baseline, on a fixed schedule or sampled traffic — conventional monitoring (status codes, latency) won’t catch a “successful but lower quality” response.

What the harness actually needs

A fixed set of representative prompts, a scoring rubric (or reference outputs), and a dashboard tracking the score trend over time — the same shape as a regression test suite, but scoring quality instead of asserting exact equality, because the output isn’t deterministic.

How do you control cost for an LLM-backed feature?

Track cost per request (not just monthly totals), cap output length explicitly, use cheaper models for lower-complexity requests via task-based routing, and cache repeated or semantically similar requests rather than regenerating them.

How do you evaluate a new model before fully switching to it?

Canary a small percentage of traffic, run the evaluation harness against both arms in parallel, compare quality, latency, and cost side by side, and ramp up gradually with a defined rollback trigger — treat it as a deploy, per this module’s Production Example.

Hands-on Lab

The gateway is this module’s routing layer — health-aware provider selection and explicit provider requests are the concrete mechanism behind “treat model choice as a routing decision.” Read the full lab documentation →

labs/async-ai-gatewayproduction-ready

The planned Model Router and Semantic Cache labs (see the Roadmap) will cover task-based routing and request caching from this module in full depth as dedicated projects.

Add cost tracking to the gateway

Wire the UsageRecord pattern from this module’s Implementation section into labs/async-ai-gateway, emitting a cost-per-request metric alongside its existing RED metrics. Write a test asserting that a request routed through a cheaper provider is reflected as a lower estimated cost than the same request routed through a more expensive one.

Version Date Change
1.0.0 2026-08-05 Initial publication.