Skip to content

Track: AI Infrastructure System Design

Listen to this page12:03
Read the transcript

1. The real exam inside the sandbox

Host: So let’s start with the thing nobody tells you walking into one of these rounds: the 45-minute system design prompt is not actually asking you to design the system. Nobody expects a correct answer for something that takes a team a year to build, and the interviewer knows that better than you do. So what’s really happening in that room?

Guest: They’re sampling five specific behaviors, and honestly the prompt is just a pretext to elicit them. Do you establish constraints before you start naming components, do you know where this class of system actually breaks, can you argue a trade-off in both directions instead of just picking a favorite, do you have any observability story at all, and do you scope honestly when you clearly can’t cover everything. That’s it — that’s the whole exam hiding inside the sandbox, and today we’re going to walk through exactly how each of those gets tested and what a strong answer sounds like versus a weak one.

2. A method built to survive interruption

Host: Okay, so if that’s the exam, what’s the actual playbook for the 45 minutes? Because I think most candidates walk in and just start drawing boxes until the clock runs out.

Guest: Right, and that’s exactly the failure mode. The method is six steps, time-boxed: constraints first for maybe five to eight minutes — who calls this, how often, how wrong can it be, whose data is it, and for AI systems specifically, what does correct even mean and how would you catch it degrading. Then you name the hard part out loud in under two minutes, something like ‘the hard part is permission filtering has to happen inside the index query, not after’ — that single sentence buys you credibility and steers everything downstream. Then you sketch one request path end to end, not a component inventory, spend your real time going deep on the piece you just named as hard, then you volunteer the failure modes and what you’d measure before anyone has to ask, and only after that do you get into scale and cost and what you’re cutting.

Host: So the ordering itself is doing work — you’re not saving the hard stuff for later, you’re front-loading it. Where does volunteering failure modes fit into that, and why does it matter so much if you do it early versus waiting to be asked?

Guest: Because the single most common way people lose this round is spending twenty minutes on boxes and then getting asked about failure modes with five minutes left, which just reads as them running out the clock before hitting the material that actually distinguishes them. If instead you volunteer it — here’s what breaks, here’s what I’d measure, before anyone asks — it signals you’ve operated one of these systems, not just designed one on a whiteboard. And practically, it means there’s still time left to actually explore it with the interviewer instead of rushing a bullet list at the buzzer.

3. Seven systems, seven dominant decisions

Host: So let’s say the prompt lands — gateway, durable agents, tool execution, RAG, serving, MCP, reliability, whatever it is. You’ve said the move is naming the dominant decision in the first five minutes. Walk me through what that actually looks like across a few of these, because I don’t think people believe there’s really just one thing per system.

Guest: There basically is, and it’s almost always a place where the obvious implementation is wrong in a specific, nameable way. Gateway: everyone reaches for per-replica rate limiting, and the dominant decision is where quota state lives — if it’s per-replica, you’ve just multiplied every tenant’s real limit by your replica count without meaning to. Durable agent execution: at-least-once is the ceiling you’re going to live under no matter what you build, so the real decision is whether your handlers are idempotent, because that’s the only thing that gets you exactly-once effect. Tool execution: scope, validation, quota, and approval feel like one blob of ‘safety,’ but they’re four separate questions and the order they run in is the design. RAG: permission filtering has to happen inside the index query itself, not as a filter you bolt on after retrieval. Serving: cold start takes minutes, so any autoscaling signal that trails load is mathematically too late — that’s not a tuning problem, it’s a structural one.

Host: So the pattern is: figure out where the naive version quietly breaks, and say that out loud before you’ve drawn a single box. What’s the payoff for doing it in minute two instead of minute twenty?

Guest: It reframes the whole conversation from ‘can you draw a coherent architecture’ to ‘do you already know where this class of system bites you’ — and that second thing is what they’re actually screening for. It also buys you the rest of the round, because the interviewer stops probing for the gotcha and starts letting you build around the constraint you already named. You end up spending forty-five minutes on the interesting trade-offs instead of forty on boxes and five getting cornered on the thing you should’ve led with.

4. Case study: the gateway’s quota trap

Host: Let’s actually run the method on one system instead of talking about it in the abstract. The async AI gateway — multiple stateless replicas behind a load balancer, and you said the dominant constraint is that nothing can be safely kept in one replica’s memory. Walk me into the trap that constraint creates.

Guest: So the natural instinct is per-tenant rate limiting, and the naive version keeps a counter in each replica’s memory — clean, fast, no external dependency. It works perfectly in a single-replica demo and then silently breaks the moment you scale to three replicas, because each replica enforces its own limit independently, so a tenant’s quota effectively multiplies by however many replicas are running. Nobody sees an error, nothing crashes, the quota just multiplies by replica count.

Host: So the fix is Redis — a shared counter every replica checks against. But you flagged that as trading one failure mode for another.

Guest: Exactly, and this is the part interviewers actually care about — what happens when Redis goes down. You have to choose deliberately: fail closed and reject requests to protect quota enforcement, which costs you availability, or fail open with a tighter local emergency limit to protect availability, which costs you weaker enforcement. Either answer is defensible, but ‘I didn’t think about it’ isn’t — because the real failure mode isn’t picking the wrong one, it’s hanging on a Redis call with no timeout and taking the whole gateway down with it.

5. Case study: the ceiling of at-least-once

Host: Let’s stay in this same failure-mode territory but move to durable agent execution — the systems that run long agent workflows with retries and checkpoints. What’s the equivalent ‘do you know the ceiling’ question there?

Guest: It’s whether a candidate accepts that at-least-once is the ceiling, full stop — exactly-once delivery isn’t on the menu. Any handler with side effects has to be idempotent, because a handler will occasionally run twice, and the classic trap is thinking a dedup key on submission solves it. It doesn’t — that prevents duplicate tasks, not duplicate deliveries of the same task, which at-least-once guarantees will happen.

Host: So where does that duplicate delivery actually bite you in practice — what’s the scenario that catches people off guard?

Guest: The zombie worker. A worker pauses long enough — GC pause, suspended VM, blocked syscall — that its lease expires, the task gets reissued and finishes elsewhere, and then the zombie wakes up still convinced it owns the work. Without a fencing token its late ack marks something done that another worker is running, or its late checkpoint overwrites newer progress — and the second question right behind it is whether you’re dead-lettering on delivery count or failure count, because a handler that reliably kills its worker never reports an explicit failure, so counting failures lets it crash-loop forever while counting deliveries bounds that at the cost of occasionally dead-lettering a task that was never actually at fault.

6. What generic systems design misses about AI

Host: Let’s zoom out from failure modes for a second, because I think there’s a bigger gap here — candidates who are genuinely good at distributed systems still stumble on these rounds. What’s actually missing?

Guest: They treat the model call as a function that returns a string. Models degrade without erroring, they cost real money per call, their latency has a long tail, and identical inputs can produce different outputs — and if your design can’t say how it detects its own quality regressing, you’ve built something nobody can operate. Cost is the other blind spot: in these systems it’s not a capacity footnote, it shapes caching, routing, context size, model selection, and if it never comes up your design reads like it’s never touched production. Then there’s retrieval — saying ‘we’ll use a vector database’ is a component choice, not a design, and the interviewer asking about RAG wants to hear about chunking, hybrid retrieval, reranking, and an actual eval set. And security has a second half nobody mentions: retrieved documents, tool descriptions, model output feeding a downstream action — perimeter auth doesn’t touch any of that.

Host: So beyond naming those gaps, is there a second thing being tested — not whether you know the right answer, but whether you can argue against yourself?

Guest: Exactly, and there are four trade-offs that come up constantly where the interviewer wants both directions, not a verdict. A gateway buys uniform policy and failover at the cost of a hop and a component to run — direct calls are simpler until a second team needs different quota rules. Shared infrastructure amortizes ops but makes isolation a property of code where a missing check becomes a breach, versus per-tenant deployment making isolation structural but growing cost linearly. Closed-by-default on policy failure protects the capability surface but turns a policy outage into a full outage, while open keeps things running but silently disables authorization — the good answer decides that per tool, not globally. And a database queue gives you transactional enqueue and kills the dual-write problem, while a broker gives you throughput at the cost of a second system to operate — say only one of those and you’ve shown a preference, not a position.

7. Turning the map into rehearsed answers

Host: So if someone’s prepping this weekend, what does ‘ready’ actually look like — not as a feeling, but as a checklist they could grade themselves against?

Guest: Seven items. You can name the dominant decision for each candidate system cold, no notes. For each system you can list three failure modes and the specific metric that would catch each one. You can argue at least four trade-offs in both directions, not just state a preference. You have a cost story, meaning you know the dominant cost line and the one lever that moves it. You have an evaluation story — how the system notices its own quality slipping, not just whether it’s up. You can say what you’d cut from v1 and defend the cut. And you’ve actually run one of the labs, so at least one answer in the room comes from something you operated, not something you read. If you want the actual question sets, every architecture page has an interview questions section written for exactly this round, and modules zero, two, four, and twelve cover decision framing, distributed systems guarantees, the AI infrastructure layer, and the observability half of every answer — that’s the map, go run it.

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

This track covers the round where you are handed a system and 45 minutes. It does not restate the questions embedded in Learn and Architecture — it says what the round is actually testing, gives a method that survives the clock, and maps each likely prompt to the page that already answers it in depth.

The prompt is a pretext. Nobody expects a correct design in 45 minutes for a system that takes a team a year. What the interviewer is sampling is narrower and more specific:

  • Do you establish constraints before proposing components? The single most reliable separator. Weaker candidates name technologies; stronger ones establish what the system must do, for whom, at what volume, and to what reliability, and let that force the components.
  • Do you know where this class of system actually breaks? Anyone can draw the happy path. Naming the failure that will actually page someone — and designing for it — is the signal.
  • Can you defend a trade-off in both directions? Not “I would use X”, but “X buys A at the cost of B, and here B is affordable because of C.” A design with no stated costs reads as a design whose costs were not understood.
  • Do you know what you would measure? A design with no observability story is one nobody has operated.
  • Do you scope honestly under time pressure? Saying “I am going to spend our time on retrieval and treat serving as a solved box, because retrieval is where the risk is” is a senior move. Trying to cover everything shallowly is not.

The Principal-level shift is from 'what would you build' to 'what would you defend'

A Staff-level answer designs a system that works. A Principal-level answer designs a system, names what it costs, identifies which decision is hardest to reverse, and says what evidence would change their mind. The last part is what interviewers remember, because it is the part that predicts how you behave when the design meets reality.

Roughly 45 minutes, with the caveat that the interviewer will interrupt and you should let them — their interruptions are the actual exam.

  1. Constraints first (5–8 min). Who calls this, how often, how fast must it answer, how wrong may it be, whose data is it, and what does it cost per call? For AI systems, add: what does “correct” mean here, and how would you detect it degrading? Write these down where both of you can see them.
  2. Name the hard part (2 min). Every one of these systems has one decision that dominates. Saying it out loud — “the hard part here is that permission filtering has to happen inside the index query, not after” — buys credibility and steers the rest of the discussion somewhere productive.
  3. Sketch the request path (8–10 min). One flow, end to end, at a level of detail you can defend. Not a component inventory.
  4. Go deep where the risk is (10–15 min). Pick the part you named as hard. This is where the round is won or lost, and where the interviewer is most likely to push.
  5. Failure modes and what you would measure (5–8 min). Volunteer these. If you are asked for them, you have already lost the point.
  6. Scale, cost, and what you cut (5 min). What breaks first at ten times the load, what the dominant cost line is, and what you deliberately left out.

Engineering Note

The most common way to lose this round is to spend twenty minutes drawing boxes and then be asked about failure modes with five minutes left. Volunteering failure modes early inverts that: it signals operational experience and it front-loads the material that distinguishes you, while there is still time to explore it.

Each of these is a plausible prompt, and each has a reference architecture with the constraints, failure modes, trade-offs, cost, and observability worked out — plus a running implementation.

Likely prompt Reference architecture Running lab
“Design an AI gateway for a company with many teams calling many providers.” Async AI Gateway async-ai-gateway
“Design a system to run long agent tasks reliably.” Durable Agent Execution durable-agent-task-engine
“How would you let agents call tools safely across an org?” Policy-Gated Tool Execution policy-gated-tool-runtime
“Design RAG over our internal documents, which have permissions.” Enterprise RAG Platform hybrid-retrieval
“Design a model serving platform for several teams.” Model Serving Platform dynamic-batching-inference
“Serve MCP to many internal teams from shared infrastructure.” Enterprise MCP Platform multi-tenant-mcp-server
“How do you know the AI platform is healthy, and what happens when it isn’t?” AI Reliability Platform slo-driven-ai-operations

Read each architecture page’s Constraints and Failure Modes sections first. They are the two that most directly convert into interview answers, and they are the two candidates most often skip.

Every prompt above has a decision that dominates the design. Naming it early is the highest-leverage thing you can do in the first five minutes.

System The decision that dominates
AI gateway Where quota state lives — per-replica buckets multiply a tenant’s real limit by the replica count
Durable agent execution At-least-once is the ceiling; exactly-once effect only comes from idempotent handlers
Tool execution Scope, validation, quota, and approval are four separate questions in a fixed order
Enterprise RAG Permission filtering happens inside the index query, never after retrieval
Model serving Cold start is minutes, so any trailing scaling signal is too late by construction
Enterprise MCP The credential must live where the SDK’s own internal calls carry it
AI reliability Alerting, scaling, and incident response must read one severity computation

Failure patterns specific to AI system design

Section titled “Failure patterns specific to AI system design”

These are the ones that show up in AI infrastructure rounds and not in generic distributed-systems rounds.

Treating the model as a black box with no failure modes

Candidates comfortable with distributed systems often design the surrounding infrastructure well and treat the model call as a function that returns a string. Models degrade without erroring, cost real money per call, have latency distributions with long tails, and return different answers to identical inputs. A design that does not account for those has not accounted for the component that motivates the system.

No evaluation story

If you cannot say how the system detects its own quality regressing, you have designed something nobody can operate. This is the single most common gap in otherwise strong AI system design answers — and it is asked about directly more often than candidates expect.

Ignoring cost per request

Traditional system design treats compute cost as a capacity-planning footnote. In AI systems, cost per request is a first-class design constraint that shapes caching, routing, context size, and model selection. A design that never mentions it reads as one that has never run in production.

Retrieval quality assumed rather than measured

“We’ll use a vector database” is a component choice, not a retrieval design. Chunking, hybrid retrieval, reranking, and a labelled evaluation set are where retrieval quality actually comes from, and the interviewer asking about RAG almost always wants that layer.

Security modelled only at the perimeter

AI systems have injection surfaces that perimeter authentication does not touch: retrieved documents, tool descriptions, and model output that feeds a downstream action. A design whose only security discussion is authentication has missed the interesting half.

Trade-offs you should be able to argue both ways

Section titled “Trade-offs you should be able to argue both ways”

The interviewer is often probing whether you hold a position or a preference. Each of these has a defensible answer in both directions, and the linked page argues both.

Build the abstraction layer, or call providers directly

A gateway buys uniform policy, failover, and cost attribution, and costs a hop plus a component to operate. Direct calls are simpler until the second team needs different quota rules. See Async AI Gateway → Trade-offs.

Shared multi-tenant service, or one deployment per tenant

Sharing amortizes operations and makes isolation a property of code — a missing check becomes a breach. Per-tenant deployment makes isolation structural and grows cost linearly. See Enterprise MCP Platform → Trade-offs.

Fail open or fail closed when the policy store is down

Closed protects the capability surface and converts a policy outage into a full outage; open keeps the system working and silently disables authorization. The good answer decides per tool. See Policy-Gated Tool Execution → Trade-offs.

Database-backed queue, or a dedicated broker

A database gives transactional enqueue alongside your own writes, removing a dual-write problem; a broker gives throughput and a second system to operate. See Durable Agent Execution → Trade-offs.

The handbook keeps interview questions next to the material that answers them, so they stay credible. For this round, the highest-yield sets are:

Before a design round

  • You can state the dominant decision for each candidate system above without looking.
  • For each system you can name three failure modes and what you would measure to catch them.
  • You can argue at least four trade-offs in both directions, not just state a preference.
  • You have a cost story: what the dominant cost line is and which design lever moves it.
  • You have an evaluation story: how the system detects its own quality regressing.
  • You can say what you would deliberately leave out of a v1, and why.
  • You have run at least one of the labs, so your answers reference something you have operated rather than something you have read.