Cheat Sheet: Design Round
Read the transcript
1. Load the sequence before you walk in
Host: So today we’re doing a cheat sheet episode, and the frame is a 45-minute design round — the kind where you get a prompt and a whiteboard and not much else. The whole point of this cheat sheet is that it’s meant to be read in the ten minutes before you walk in, not consulted during, because the value isn’t the content, it’s having the order already loaded into your head. So let’s walk through that sequence minute by minute.
Guest: Right, and the sequence is pretty rigid on purpose. First eight minutes are constraints — volume, p99, correctness, data boundary, cost per call, failure budget — you say out loud who calls this, how often, how fast, how wrong it’s allowed to be, whose data it touches. Then minutes eight to ten you name the one hard part, because every one of these systems has exactly one decision that dominates everything else, and saying that out loud tells the interviewer you know where the weight is. From there it’s tracing one request end to end at a defensible depth — not a component inventory — then twenty to thirty-three is going deep on that hard part, which is where the round is actually won, then you volunteer failure modes and what you’d measure, and close with scale, cost, and what you’d cut at ten x. And the thing to internalize now, before you’re in the room, is that interruptions aren’t a disruption to this sequence — they are the actual exam.
2. The arithmetic and the one hard part per system
Host: Okay, before we get to the one hard part per system, let’s arm people with the actual math, because ‘we’ll need a few replicas’ is not an answer an interviewer accepts. What’s the formula you want loaded and ready?
Guest: Little’s Law — concurrency equals arrival rate times average latency. So 200 requests per second at 2 seconds average latency means 400 requests in flight at any moment, and that single number sets your pool sizes, your semaphore bounds, your replica count. Size for peak plus roughly thirty percent headroom, say it out loud, and you’ve just turned a vibe into an engineering decision.
Host: That’s the number, now give me the one dominant decision per system, fast, because that’s what minutes eight to ten are for.
Guest: Gateway — where quota state lives, because per-replica buckets silently multiply a tenant’s limit by replica count. Durable execution — at-least-once is the ceiling, so exactly-once effect only comes from idempotent handlers. Tool execution — scope, validation, quota, approval, in that fixed order. RAG — permission filtering has to happen inside the index query, never bolted on after retrieval. Serving — cold start takes minutes, so any trailing scaling signal is already too late by construction. MCP — the credential has to live wherever the SDK’s own internal calls carry it. Reliability — alerting, scaling, and incident response all have to read one shared severity computation. And every trade-off you cite — gateway versus direct calls, shared versus per-tenant, fail open versus closed, database queue versus broker — you argue both directions out loud, because picking a side without acknowledging the cost of that side is the tell that you’ve memorized an answer instead of reasoned to it.
3. The red flags that sink an otherwise strong answer
Host: So let’s close with the tells — the things that sink an answer even when the boxes and arrows look right. What’s the first thing that makes you wince as an interviewer?
Guest: Naming technologies before you’ve said what constraint forces that choice — that’s the first one, and adjectives are the second, ‘scalable’ and ‘high throughput’ with no arithmetic behind them. Then treating the model as a pure function with no cost, no latency tail, no degradation, and having no evaluation story at all, which is the single most common gap even in strong answers. Add security reduced to just authentication while retrieved documents and tool descriptions and model output feeding an action go unmentioned, spending twenty minutes drawing boxes and reaching failure modes with five minutes left, and a design with no stated cost anywhere — because that reads as a design whose costs were never understood, and if you run that list against your own answer in real time, you’ll catch yourself before the interviewer has to.
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Use This When
Section titled “Use This When”You have 45 minutes and a prompt. Read it in the ten minutes before the round, not during — the value is in the order, and the order only helps if it is already loaded.
Depth lives in Track: AI Infrastructure System Design and Module 13.
The Sequence
Section titled “The Sequence”| Min | Do | Say out loud |
|---|---|---|
| 0–8 | Constraints. Volume, p99, correctness, data boundary, cost per call, failure budget | “Before components — who calls this, how often, how fast, how wrong may it be, whose data?” |
| 8–10 | Name the hard part. One decision dominates every one of these systems | “The hard part here is X, so that’s where I’ll spend our time” |
| 10–20 | One request path, end to end, at a defensible depth. Not a component inventory | “Let me trace one request, then go deep where the risk is” |
| 20–33 | Go deep on the hard part. This is where the round is won | — |
| 33–40 | Failure modes and measurement. Volunteer these | “Here’s how this breaks, and what I’d watch” |
| 40–45 | Scale, cost, what you cut | “At 10× the first thing to break is X. I’d leave Y out of v1 because Z” |
Let interruptions happen. They are the actual exam.
Quick Reference
Section titled “Quick Reference”Little’s Law — concurrency = arrival_rate × average_latency. 200 rps × 2 s = 400 in flight.
That number sets pool sizes, semaphore bounds, and replica count. Size for peak plus ~30% headroom.
The one hard part, per system
| System | The decision that dominates |
|---|---|
| AI gateway | Where quota state lives — per-replica buckets multiply a tenant’s limit by replica count |
| Durable execution | At-least-once is the ceiling; exactly-once effect only from idempotent handlers |
| Tool execution | Scope, validation, quota, approval — four questions in a fixed order |
| Enterprise RAG | Permission filtering inside the index query, never after retrieval |
| Model serving | Cold start is minutes, so any trailing scaling signal is too late by construction |
| Enterprise MCP | The credential must live where the SDK’s own internal calls carry it |
| Reliability | Alerting, scaling, and incident response read one severity computation |
Trade-offs, both directions — gateway vs direct calls · shared vs per-tenant · fail open vs
closed · database queue vs broker. Each is argued out on the matching
Architecture page under #trade-offs.
Red Flags
Section titled “Red Flags”- Naming technologies before establishing a single constraint.
- Capacity in adjectives — “high throughput”, “scalable” — with no arithmetic.
- Treating the model as a function returning a string: no cost, no latency tail, no degradation.
- No evaluation story. The most common gap in otherwise strong answers, and asked about directly.
- Security discussed only as authentication, ignoring retrieved documents, tool descriptions, and model output feeding an action.
- Twenty minutes of boxes, then failure modes with five minutes left.
- A design with no stated cost — it reads as one whose costs were never understood.