Module 13: System Design
Read the transcript
1. The skill this module actually teaches
Host: So we’ve spent other modules going layer by layer — storage, networking, the model itself. This module is different: it’s not another layer, it’s the thing that spans all of them. What’s the actual skill we’re naming here?
Guest: Right, it’s not knowing more components, which is the trap people fall into when they hear ‘system design.’ The skill is turning a vague request into something with known behavior, known cost, and a known failure envelope before you write a line of code. That means establishing constraints before you propose solutions, modelling capacity in numbers instead of adjectives like ‘fast’ or ‘scalable,’ identifying the one decision that’s hardest to undo, and stating up front what evidence would prove your design wrong.
Host: And you’re saying AI components turn up the pressure on every one of those four things at once.
Guest: Exactly — the model is the piece that usually motivates building the system, and it breaks the normal rules. It degrades quietly instead of throwing errors, it costs real money on every single call, and it can give you a different answer to the identical input twice in a row. So the constraints, the numbers, the irreversibility, the falsifiability — none of that is optional anymore, and none of it is forgiving if you skip it.
2. A design is a set of falsifiable claims, not a diagram
Host: So if I hand you a whiteboard with boxes and arrows, you’re telling me that’s not actually a design yet.
Guest: Right, it’s a diagram, and a diagram just says what the parts are. A design says what the system does — this many requests per second, at this p99, at this cost per call, degrading this specific way when the index goes stale. Every one of those is a claim measurement can prove wrong, and if a claim can’t be refuted, it’s decoration, not design.
Host: Okay, so walk me through why that framing actually changes what you do first. What falls out of treating a design as a set of claims?
Guest: Three things. First, constraints come before components — pick the database before you know the constraint it satisfies and you’ve made a preference, not a decision, and you’ve got nothing to defend it with later. Second, reversibility is the axis that matters — most choices can be undone in an afternoon and deserve an afternoon, but a few, like the data boundary or the external contract, are expensive for years, so that’s where the real thinking time goes. And third, if you can’t measure a claim you can’t operate the system, which for AI means a quality signal alongside the latency one, because fast, cheap, and quietly wrong is exactly the failure this discipline is built to catch.
3. The method as a loop, and where the expensive mistakes concentrate
Host: So walk me through this as an actual flow, not just a checklist. You said the design is a set of claims — where does the loop close?
Guest: You draft the claims, you draft the architecture that would satisfy them, and then you go measure whether it actually does. If the numbers disagree with the claim, you don’t patch the diagram, you go back and revise the claim or the design that produced it. The loop is the whole method — a design isn’t done when it’s drawn, it’s done when measurement agrees with it, and until then you’re still designing whether you feel like it or not.
Host: And that’s where the ‘hardest to reverse’ branches come in — you named data boundary and external contract already. What’s the third, and why do these three specifically eat the budget?
Guest: State ownership — what a failure loses. If you don’t decide up front who owns the source of truth for a piece of state, a crash doesn’t just cost you uptime, it costs you data you can’t reconstruct. Data boundary decides who can see what and you basically cannot retrofit it once other systems have grown around your leaky version. External contract decides how long you’re stuck carrying a design after you’ve outgrown it. None of these are exhaustive, but in AI systems that’s where the expensive mistakes concentrate, so that’s where the thinking time actually goes.
4. Six questions that make constraints numbers, not adjectives
Host: So once you know where the expensive mistakes live, how do you actually pin the constraints down before you touch a component? You said earlier a design has to be falsifiable — where does that start?
Guest: It starts with six questions, and the rule is the answer has to be a number or a named party, not an adjective. Who calls this and how often — requests per second at peak, not average, because a batch job at 9am and even traffic across the day can hit the same daily total with wildly different capacity needs. How fast must it answer — a p99, not a mean, plus what happens when you miss it: does the caller retry, queue, or just fail.
Host: Those two feel like standard systems questions. Where does it start diverging for AI specifically?
Guest: Right at question three — how wrong may it be, and how would you know. ‘Correct’ needs a definition, and that definition needs an eval set, or you have literally no quality signal, just vibes. Then whose data is it — tenant, region, retention, the constraint that’s discovered latest and costs the most to fix afterward — what does a call cost, tokens in, tokens out, retrieval hops, retries, because in AI systems cost per request is a design constraint, not a footnote — and what’s the failure budget, which is the input to every redundancy decision downstream.
Host: And once you’ve got those six pinned down, you still don’t know how the system actually behaves under load — that’s the workload shape, arrival pattern, payload sizes, that kind of thing?
Guest: Exactly, and AI workloads have two properties generic services don’t. Request cost swings by an order of magnitude depending on context length. And latency is dominated by a component whose behavior you don’t control.
5. Little’s Law and the arithmetic nobody does
Host: So let’s actually do the arithmetic instead of gesturing at it. If I know arrival rate and latency, what do I get for free?
Guest: Concurrency, exactly. Little’s Law: concurrency equals arrival rate times average latency. Two hundred requests per second at two seconds average latency means four hundred requests in flight at any given moment, and that number is derivable before a single line of code exists.
Host: And that number isn’t academic — it’s what sizes your connection pool, your replica count, your memory footprint.
Guest: Right, which is why the Workload and Capacity code is worth looking at directly — it takes peak rps, average latency, and token costs, runs the same Little’s Law arithmetic, and sizes replicas off usable concurrency per replica after headroom. Drop headroom from thirty percent to ten and you cut replica count by roughly a fifth, but you’ve also pushed the system right up against the point where queue-wait time becomes your latency instead of the model call.
Host: So the model isn’t a forecast, it’s a set of numbers to argue with — if it says four hundred concurrent and your pool is configured for a hundred, you’ve found a bug before it shipped.
Guest: That’s the whole payoff. And the trick that makes it actually useful: run it twice, once at the given constraints and once at ten times the request rate. Whatever breaks first in that second run is the single most important fact about the design, and it’s also the exact question interviewers ask most reliably.
6. Finding the decision that is hardest to undo
Host: So before you even get to stress-testing a design, there’s a prior question: which decisions in this design are even worth agonizing over? Because I don’t think every choice deserves the same amount of deliberation.
Guest: Right, and that’s the organizing question underneath the whole method: for each choice, if you’re wrong, what does fixing it cost? A retry policy is a config change, so let whoever’s closest to it just decide and move on. But a data model that mixes two tenants’ embeddings into one index — that’s not a fix, that’s a migration, an audit, and a disclosure. Most decisions are two-way doors, walk back through them cheaply. A few are one-way doors, and those are the ones that deserve the meeting.
Guest: There’s actually a caveat worth adding here, since the failure runs both directions. Treating every decision like a one-way door is the common failure — it’s slow, it breeds committees, nothing ships. But the rarer failure is worse: treating an actual one-way door as if it were reversible, shipping the tenant-mixed index because it was Tuesday and someone wanted to move fast. The whole skill is classification before commitment, not caution as a default setting.
7. Case study: async-ai-gateway’s constraints made concrete
Host: Let’s ground that classification instinct in something concrete. The async-ai-gateway lab keeps coming up — walk me through how one of its constraints actually turned into a design decision, not just a feature.
Guest: Take ‘many teams, each with their own quota.’ That’s not a feature request, it’s a constraint, and it forces per-tenant rate limiting instead of one global limit. But it immediately raises the one-way-door question: where does the quota counter live? If it’s in-process memory, it’s correct on one replica and silently wrong on three — each replica enforces the limit independently, so a tenant gets three times their quota and nobody notices until they’ve built traffic patterns against the wrong number.
Host: So the fix has to be atomic across replicas, not just correct on a laptop. What does the lab actually do about that, and how would you know it’s true rather than assumed?
Guest: It uses a Redis-backed token bucket where refill and consume happen in one Lua script — atomic by construction, not by convention. And it’s checkable: at a capacity of 20 with two replicas, the atomic version allows 20 total, the naive get-then-set version allows all 40. The other two constraints follow the same pattern. ‘A provider outage must not become our outage’ produces health-aware routing and a circuit breaker, with the claim that a failing provider gets ejected within a bounded number of failures — verified in the test suite, not asserted in a design doc. And ‘a rolling deploy must not kill in-flight requests’ produces explicit draining with a stated timeout: in-flight work finishes within 30 seconds or gets abandoned on purpose, which is different from being killed by accident.
Host: So all three constraints are recoverable straight from the code — the quota script, the ejection test, the drain timeout — because someone wrote the claim down before writing the component.
Guest: Exactly, and that’s the whole payoff. A design whose constraints live in someone’s memory turns into a system nobody will touch within two quarters — not because the code is bad, but because nobody can tell which behavior is load-bearing anymore. Here, if you want to know what the system promises, you don’t ask the person who built it, you read the Lua script, the breaker test, and the drain timeout.
8. Where designs quietly fail
Host: So let’s talk about how these designs actually fall apart in practice, because I suspect it’s the same handful of ways every time. What’s the most common one you see, the one that’s obvious even from outside the project?
Guest: A design that opens with a technology list. The moment someone starts naming the specific tools before anything else, you know they skipped the step that would justify those choices. The tell is simple: if you changed the requirements, would the diagram change? If not, the diagram was never derived from anything, and it won’t survive the first hard question. Close behind that is the adjective problem — those aren’t constraints, they’re vibes. Nobody notices until launch, when the connection pool or queue bound turns out to have been picked by whatever the library defaulted to.
Host: Those two feel almost like bad habits you could catch with a checklist. But you’ve said the AI-specific failure is different — it doesn’t even show up as a failure. Walk me through that one.
Guest: Right, this is the one that should keep people up at night. The system is up, it’s fast, it’s within budget — every dashboard is green — and the answers have been quietly getting worse for six weeks because an index rebuild changed how documents got chunked. Latency monitoring is structurally blind to this; it’s not a latency problem or an error-rate problem, it’s a quality problem, and only a scheduled eval set catches it. And that’s really one instance of a bigger pattern: cost blindness, where nobody priced retries or ballooning context until the bill explained it, and irreversibility, where tenant data lands in a shared index or state gets owned by two components at once, cheap that week and expensive for years. Then there’s the mirror image — building for a hundred times your actual traffic and paying the operational tax on that complexity forever. Same discipline fixes all of it: state the constraint, and let it justify the complexity or refuse to.
9. The trade-offs that don’t resolve cleanly
Host: So none of these trade-offs actually resolve into a clean rule. Let’s start with up-front versus iterative design, since that feels like the oldest argument in the room.
Guest: It never resolves cleanly because both sides are true. Up-front design catches the one-way doors while they’re cheap, but it costs time before anyone’s served; iterative design finds real requirements faster but quietly accumulates decisions you can’t undo. The split that actually works follows reversibility — decide the data boundary, state ownership, the external contract deliberately and early, and let everything else be discovered by building.
Host: That same shape shows up in platform versus specific-case, doesn’t it — build the general thing or just ship the one thing in front of you?
Guest: Exactly, and the tell is the second consumer, not the first. One use case is just a use case; a platform pays for itself around the third consumer, so before that it’s speculative generality you’re operating for nothing. Same logic covers abstraction — buy flexibility only where you can name the specific change you expect, commit everywhere else — and at Principal scope, a consistent stack usually beats the better-fitting fifth datastore once you count what it costs to run.
10. Security, performance, and scaling as one AI-shaped problem
Host: Let’s pull three things together that usually get separate chapters — security, performance, and scaling — because in these systems they’re really the same constraint viewed from different angles. Start with security, since you said the data boundary decision from earlier is really a security decision too.
Guest: Right, deciding which tenants share storage or a cache isn’t just a scaling question, it’s the whole security posture. Isolation by code — trusting every query path to remember the filter — is weaker than isolation by structure, and if you go shared, which is usually correct, the filter has to sit inside the query at a chokepoint everything crosses, not bolted on after retrieval. And AI adds surfaces perimeter auth never covers: retrieved documents, tool descriptions, model output feeding an action — a design whose security section is just authentication has handled the smaller half of the problem.
Host: So that same boundary decision now has to answer an audit question too — who accessed what, on whose behalf?
Guest: Exactly, and that’s cheap if identity flows through the system from the start, nearly impossible to bolt on later. Now performance runs into the identical structure — you model the p99 as a sum across admission, retrieval, model call, post-processing, and design against the tail, because fan-out means a ten-backend request sees roughly the p99 of one of them as its typical case, not the mean.
Host: And the thing that actually saturates first isn’t CPU, which is what trips people up when they go to scale this.
Guest: Right, it’s provider concurrency or accelerator memory, so autoscaling on CPU is measuring the wrong thing entirely — you need queue wait measured separately from execution, or a starved system looks identical to a genuinely slow one on the chart. And scaling closes the loop: shard along the same boundary you chose for security, scale request handling and inference independently since their curves are unrelated, and remember cost scales with traffic here, so ten times the load is roughly ten times the model spend, which is what actually decides how far this design goes before it needs a different shape.
11. Saying it out loud: how this plays in an interview
Host: So say someone asks you this cold in an interview — walk me through your opening move, the thing you say in the first five minutes that signals you actually know what you’re doing.
Guest: Constraints before components, always. Then you name the decision that’s hardest to reverse — naming that hard part explicitly is the single highest-leverage thing you can do.
Host: And when the conversation turns into disagreement — which it will, someone pushes back on your design — how do you keep that from turning into a shouting match about taste?
Guest: You separate it into three buckets: constraint, evidence, or preference. If you disagree about constraints, that’s actually an unstated requirement, resolve that first. If constraints agree but conclusions differ, name the evidence that would settle it and what it costs to get. If it’s genuinely preference within the same constraints, it’s a two-way door and not worth arguing over — and if they ask for a worked example, you say something like: hold p99 under 800 milliseconds at 200 requests per second, that’s 400 in flight by Little’s Law so the pool is sized for 400 with headroom, and if p99 blows past that at lower load, the model’s wrong and you check queue wait against provider latency first because those need opposite fixes — plus recall@10 stays above 0.85 or retrieval has regressed, checked on every index rebuild since that’s the event that causes it, not on a schedule.
12. Practicing the method, and what to read next
Host: So if someone wants to actually drill this rather than just nod along, where do they start?
Guest: Reverse it on something real: take async-ai-gateway and try to recover its constraints just from the code — what must have been true for per-tenant limiting, circuit breaking, and explicit draining to be worth the engineering effort? Write your guesses down before you look, then compare against the architecture page that actually states them. The gap between what you inferred and what was intended is the part of design that never survives into code, which is exactly the argument for writing it down in the first place. Then do the uncomfortable version on your own system: fill in real numbers, check the computed concurrency against your actual pool config and the computed cost against your actual invoice — in most systems one of those is off by more than double, and that gap is a finding either way.
Host: That’s a good place to leave people. Anything you’d point them to afterward, for where these ideas actually come from?
Guest: Four things, in order of how often you’ll use them: Bezos’s one-way and two-way door framing from the 2015 shareholder letter, because it’s the most portable idea in the whole module. Kleppmann’s Designing Data-Intensive Applications for the storage and consistency decisions this method routes you into. The SRE book’s chapters on SLOs and error budgets, for how a design claim becomes an operational commitment somebody’s on call for. And Ousterhout’s A Philosophy of Software Design for when an abstraction is actually earning its cost. And this handbook’s own ADRs are worth reading cold, since they’re worked examples of the written output this whole module is arguing you should produce.
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Executive Summary
Section titled “Executive Summary”The other modules in this track each cover one layer. This one covers the activity that spans them: turning an underspecified request into a system whose behavior, cost, and failure envelope are known before anyone writes the first line.
The distinguishing skill is not knowing more components. It is establishing constraints before proposing any, modelling capacity and cost in numbers rather than adjectives, identifying which decision is hardest to undo, and stating what evidence would prove the design wrong. AI systems raise the stakes on all four, because the component that motivates the system — the model — degrades without erroring, costs real money per call, and returns different answers to identical inputs.
Mental Model
Section titled “Mental Model”A design is a set of falsifiable claims, not a diagram.
A diagram says what the parts are. A design says what the system will do — this many requests per second, at this p99, for this cost per call, degrading in this specific way when the retrieval index is stale — and each of those is a claim that measurement can refute. Anything that cannot be refuted is decoration.
Three consequences follow, and they organize the rest of this module:
- Constraints come before components. A component chosen before the constraint it satisfies is a preference, not a decision. The constraint is what makes the choice defensible later, and what tells you when it stops being right.
- Reversibility is the axis that matters. Most decisions can be undone in an afternoon and deserve an afternoon of thought. A few — the data boundary, who owns which state, the external contract — are expensive for years. Spend your design time in proportion.
- A design you cannot measure is a design you cannot operate. Every claim needs a signal, and for AI systems that means a quality signal as well as a latency one. A system that is fast, cheap, and quietly wrong is the failure mode this discipline exists to prevent.
The output of a design is a decision record, not a picture
What survives contact with a team is not the diagram — it is the written record of what was decided, what it cost, what was considered and rejected, and what would change the answer. Six months later nobody remembers why the queue is in the database instead of a broker, and the absence of that record is what causes the rewrite. The ADR section exists for exactly this reason.
Architecture
Section titled “Architecture”The method, as a flow. The loop at the bottom is the point: a design is not finished when it is drawn, it is finished when measurement agrees with it.
flowchart TD
START([New AI system]) --> CONSTRAINTS[Establish constraints<br/>volume, latency, correctness,<br/>data boundary, cost per call]
CONSTRAINTS --> WORKLOAD[Characterize the workload<br/>arrival shape, payload size,<br/>read/write mix, tail sensitivity]
WORKLOAD --> CAPACITY[Model capacity and cost<br/>Little's Law, tokens per call,<br/>dominant cost line]
CAPACITY --> HARD{Which decision is<br/>hardest to reverse?}
HARD --> DATA[Data boundary<br/>and permission model]
HARD --> STATE[State ownership<br/>and durability]
HARD --> CONTRACT[External contract<br/>and versioning]
DATA --> DECOMPOSE[Decompose into components<br/>with one owner each]
STATE --> DECOMPOSE
CONTRACT --> DECOMPOSE
DECOMPOSE --> EVAL[Define the quality signal<br/>eval set, thresholds,<br/>regression gate]
EVAL --> OBSERVE[Define the operational signal<br/>SLOs, burn rate, cost per request]
OBSERVE --> REVERSIBLE{Reversible<br/>within one deploy?}
REVERSIBLE -->|Yes| SHIP[Ship behind a flag<br/>and measure]
REVERSIBLE -->|No| EVIDENCE[Name the evidence that<br/>would change the decision]
EVIDENCE --> SHIP
SHIP --> MEASURE[Compare against the<br/>declared thresholds]
MEASURE -->|Within budget| DONE([Design holds])
MEASURE -->|Outside budget| CONSTRAINTSThe three branches out of “hardest to reverse” are not exhaustive, but in AI systems they are where the expensive mistakes concentrate. The data boundary determines who can see what and is nearly impossible to retrofit. State ownership determines what a failure loses. The external contract determines how long you carry the design after you have outgrown it.
Deep Dive
Section titled “Deep Dive”Establishing constraints. Six questions, and the answers should be numbers or named parties:
- Who calls this, and how often? Requests per second at peak, not average — and the shape, because a system serving a batch job at 09:00 has different capacity needs than one serving even traffic at the same daily total.
- How fast must it answer? A p99 target, not a mean, and a statement of what happens when it is missed: does the caller retry, queue, or fail?
- How wrong may it be, and how would you know? The AI-specific one. “Correct” needs a definition, and the definition needs an eval set, or the system has no quality signal at all.
- Whose data is it? Tenant, region, and retention. This is the constraint most often discovered late and most expensive to satisfy afterwards.
- What does a call cost? Tokens in, tokens out, retrieval hops, retries. Cost per request is a design constraint in AI systems, not a capacity-planning footnote.
- What is the failure budget? How much unavailability is acceptable, which is the input to every redundancy decision downstream.
Characterizing the workload. Constraints tell you the target; workload shape tells you what the system will actually experience. Arrival pattern (smooth, bursty, scheduled), payload size distribution, read/write mix, and tail sensitivity. AI workloads have two properties that generic services do not: request cost varies by an order of magnitude with context length, and latency is dominated by a component whose behavior you do not control.
Modelling capacity. Little’s Law is the whole of it for a first pass:
concurrency = arrival_rate × average_latency
At 200 requests per second with an average latency of 2 seconds, 400 requests are in flight at any moment. That number determines connection pool sizes, semaphore bounds, memory, and how many replicas you need — and it is derivable before any code exists. Most capacity surprises are this arithmetic not being done.
Finding the irreversible decision. Ask of each choice: if this is wrong, what does fixing it cost? A retry policy is a config change. A queue technology is a week. A data model that mixes two tenants’ embeddings in one index is a migration, an audit, and a disclosure. Rank by that, then spend design effort in the same order.
Research Note
The framing that most reliably improves design discussions: most decisions are two-way doors and should be made fast by the people closest to them; a few are one-way doors and deserve deliberation. The failure mode in engineering organizations is treating every decision as a one-way door, which is slow, and the rarer, worse one is treating a one-way door as reversible.
Source: Jeff Bezos, 2015 Amazon shareholder letter (one-way and two-way doors)
Making it falsifiable. For each claim in the design, name the measurement that would refute it and the threshold at which you would act. “Retrieval quality is sufficient” is not a claim. “Recall@10 stays above 0.85 on the labelled set, checked on every index rebuild” is.
Implementation
Section titled “Implementation”Capacity modelling is the part of design most often waved at, and it is straightforwardly executable. This is the arithmetic above, written so it can be run against real constraints instead of estimated in a meeting:
A first-pass capacity and cost model
from dataclasses import dataclass
@dataclass(frozen=True)class Workload: """Constraints, stated as numbers so the design can be checked against them."""
peak_rps: float avg_latency_seconds: float p99_latency_target_seconds: float input_tokens: int output_tokens: int input_cost_per_1k: float output_cost_per_1k: float
@dataclass(frozen=True)class Capacity: concurrent_requests: float replicas_needed: int cost_per_request: float cost_per_day: float
def model(workload: Workload, *, concurrency_per_replica: int, headroom: float = 0.3) -> Capacity: # Little's Law: everything in flight at a given instant. in_flight = workload.peak_rps * workload.avg_latency_seconds
# Size for peak plus headroom, because a replica running at its concurrency # limit is a replica whose queue-wait time is about to become the latency. usable = concurrency_per_replica * (1 - headroom) replicas = max(1, ceil_div(in_flight, usable))
cost = ( workload.input_tokens / 1000 * workload.input_cost_per_1k + workload.output_tokens / 1000 * workload.output_cost_per_1k ) return Capacity( concurrent_requests=in_flight, replicas_needed=replicas, cost_per_request=cost, cost_per_day=cost * workload.peak_rps * 86_400, )
def ceil_div(numerator: float, denominator: float) -> int: return int(-(-numerator // denominator))Two things this makes visible that prose does not. First, headroom is a design decision with a
cost — dropping it from 0.3 to 0.1 cuts the replica count by roughly a fifth and moves the system
much closer to the point where queue-wait dominates latency. Second, cost_per_day computed at peak
is deliberately pessimistic; the honest version needs the arrival shape, and discovering that you
cannot supply it is itself a finding.
The output is not a plan, it is a set of numbers to argue with. If the model says 400 concurrent requests and the connection pool is configured for 100, that is a bug found before it was written.
Engineering Note
Run this model twice: once with the constraints as given, and once at ten times the request rate. The second run tells you which component breaks first, which is the single most useful thing to know about a design and the question interviewers ask most reliably.
Production Example
Section titled “Production Example”The async-ai-gateway lab is a worked instance of this method,
and its structure is traceable back to specific constraints.
The constraint “many teams, each with their own quota” produced per-tenant rate limiting rather than a global limit — and immediately produced the irreversible-decision question: where does quota state live? In-process is correct on one replica and wrong on three, where each replica enforces the quota independently and a tenant silently receives three times their limit. That is a one-way door in the sense that matters: by the time you notice, tenants have built against the wrong limit.
The constraint “a provider outage must not become our outage” produced health-aware routing and a circuit breaker, and the falsifiable claim that a failing provider is ejected within a bounded number of failures. The lab’s test suite is where that claim is checked, which is the difference between a design and an intention.
The constraint “a rolling deploy must not kill in-flight requests” produced explicit draining, with a stated timeout — a claim that in-flight work completes within 30 seconds or is abandoned deliberately rather than by accident.
The constraint you did not write down is the one that changes
Every one of the above is recoverable from the code, but only because it was written down first. A design whose constraints live in someone’s memory becomes, within two quarters, a system nobody is willing to change — not because the code is bad, but because nobody can tell which behavior is load-bearing.
Failure Modes
Section titled “Failure Modes”Components chosen before constraints
The most common failure, and the one most visible from outside. A design that opens with a technology list has skipped the step that makes the list defensible, and it will not survive the first question about why. The tell is a design where changing the requirements would not change the diagram.
Capacity estimated in adjectives
“High throughput”, “low latency”, and “scalable” are not constraints, and a design built on them cannot be checked. The failure surfaces at launch, when the connection pool, the queue bound, or the replica count turns out to have been chosen by default rather than by arithmetic.
No quality signal
The AI-specific failure. The system is available, fast, and within budget, and its answers have been degrading for six weeks because an index rebuild changed the chunking. Latency monitoring cannot see this. Only an eval set run on a schedule can, and a design without one has no way to detect its own most likely failure.
Cost discovered in the invoice
Cost per request is a design constraint that shapes caching, context size, model selection, and routing. Systems designed without it typically discover at scale that the dominant cost line is something nobody considered — retries, or context that grew as features were added, or a retrieval step firing on every request including the ones that did not need it.
Irreversible decisions made casually
Tenant data landing in a shared index, a public API shipped without a version, state owned by two components at once. Each is cheap in the week it happens and expensive for years. The absence of a moment where someone asked “how hard is this to undo?” is the root cause every time.
Designing for a scale that never arrives
The mirror image, and it costs just as much. A system built for a hundred times its actual traffic carries the operational burden of that complexity every day, in exchange for a capacity nobody uses. The discipline is the same: state the constraint, and let it justify the complexity or not.
Trade-offs
Section titled “Trade-offs”Design up front, or design as you go
Up-front design catches the irreversible mistakes while they are still cheap and costs time before any user is served. Iterative design finds real requirements faster and accumulates decisions that are hard to undo. The workable split follows reversibility: decide the one-way doors — data boundary, state ownership, external contract — deliberately and early, and let everything else be discovered.
Build the platform, or solve the case in front of you
A platform amortizes across teams and pays for itself at the third consumer; before that it is speculative generality with an operational cost. Solving the specific case ships sooner and risks three divergent implementations that must later be reconciled. The signal to watch is the second consumer, not the first — one is a use case, two is a pattern.
Optimize for change, or optimize for the known workload
Abstraction layers, feature flags, and pluggable interfaces buy the ability to change direction and cost indirection that every future reader pays for. Committing to the known workload is faster and simpler until the workload moves. Buy flexibility only where you can name the change you expect, and commit everywhere else.
Consistency across systems, or the right tool per system
A consistent stack means one set of runbooks, one on-call rotation, and engineers who can move between services. Per-system choices fit each problem better and multiply the operational surface. At Principal scope the second-order cost usually dominates: the fifth datastore is rarely worth what it costs to operate, even when it is genuinely the better fit.
Security
Section titled “Security”- The data boundary is a design decision, not an implementation detail. Which tenants, regions, and classifications share storage, an index, or a cache is decided at design time and is extremely expensive to change afterwards. Decide it explicitly and write it down.
- Isolation by code is weaker than isolation by structure. In a shared multi-tenant system a missing filter is a breach; in a per-tenant deployment it is a bug. Choosing shared is usually right, and it means the filter belongs at a chokepoint that every path must cross, with a test that proves it.
- Permission filtering belongs inside the query, not after it. Retrieving then filtering leaks through result counts, latency, and any path that forgets the second step. This is the design-time form of the failure that the Enterprise RAG Platform page treats in depth.
- AI systems have injection surfaces that perimeter auth does not cover — retrieved documents, tool descriptions, and model output that feeds a downstream action. A design whose only security section is authentication has addressed the smaller half.
- Design for the audit question up front. “Who accessed what, when, and on whose behalf” is cheap to answer if identity flows through the system and nearly impossible to reconstruct afterwards.
Performance
Section titled “Performance”- Model the latency budget as a sum, not a target. Allocate the p99 across the path — admission, retrieval, model call, post-processing — and the component with no room becomes obvious before it is built.
- Design against the tail. Mean latency is a property of the system nobody experiences. Fan-out makes this worse arithmetically: a request touching ten backends sees roughly the p99 of one of them as its typical case.
- Know which resource saturates first. For AI systems it is usually provider concurrency or accelerator memory, not CPU — which is why autoscaling on CPU is the wrong signal for most of these systems.
- Measure queue wait separately from execution. A system starved of concurrency slots and a system that is genuinely slow look identical on an end-to-end latency chart and need opposite fixes.
- Cost and latency trade against each other explicitly. Caching, smaller models, and shorter context all move both numbers, usually in opposite directions. Decide which one the design is optimizing and say so.
Scaling
Section titled “Scaling”- Ask what breaks first at ten times, not what the architecture supports. The answer is almost always a specific bounded resource — a connection pool, a single-writer database, a provider quota, an in-process cache — and naming it is more useful than any capacity claim.
- Statelessness is what makes horizontal scaling work, so every piece of in-process state is a scaling decision: rate limiter buckets, caches, dedup sets, and session data each need an answer for what happens on the second replica.
- Shard along the isolation boundary you already chose. If tenancy is the data boundary, it is usually also the right partition key, and using two different boundaries doubles the coordination.
- Scale the components independently. Request handling, embedding generation, and batch inference have unrelated scaling curves and should not share a deployment unit.
- Cost scales with traffic in AI systems, unlike most infrastructure, where fixed capacity absorbs growth. Ten times the traffic is roughly ten times the model spend, which makes cost per request the constraint that governs how far the design can scale before it must change shape.
Interview Questions
Section titled “Interview Questions”How do you start a system design you have never seen before?
Constraints before components: who calls it and how often, how fast it must answer, how wrong it may be and how you would detect that, whose data it is, and what a call costs. Then characterize the workload shape, do the Little’s Law arithmetic, and name the decision that is hardest to reverse. Naming the hard part out loud in the first five minutes is the highest-leverage move available.
How do you decide how much design to do before building?
By reversibility. Decisions that can be undone in a deploy get made quickly by whoever is closest to them. Decisions that are expensive for years — the data boundary, state ownership, the external contract — get deliberated. The organizational failure mode is treating everything as a one-way door, which is slow; the rarer and worse one is treating a one-way door as reversible.
What makes designing AI systems different from designing conventional distributed systems?
Four things. The dominant component degrades without erroring, so quality needs its own signal and its own regression gate. Cost per request is a first-class design constraint rather than a capacity footnote. Latency is dominated by something you do not control and cannot profile. And the system has injection surfaces — retrieved documents, tool descriptions, model output feeding an action — that perimeter authentication does not touch.
Your design is challenged by a senior colleague who prefers a different approach. What do you do?
Separate the disagreement into constraint, evidence, and preference. If you disagree about constraints, resolve that first — most design arguments are actually about unstated requirements. If the constraints agree and the conclusions differ, name the evidence that would settle it and what it would cost to get. If it is genuinely preference within the same constraints, it is a two-way door and not worth the meeting.
Making a design falsifiable, concretely
“The claim is that we hold p99 under 800 ms at 200 requests per second. That is 400 requests in flight by Little’s Law, so the pool is sized for 400 with 30% headroom. If p99 exceeds 800 ms at under 200 rps, the model is wrong and the first thing I would check is queue wait versus provider latency, because those need opposite fixes. Separately, recall@10 on the labelled set stays above 0.85 or retrieval has regressed, and that runs on every index rebuild rather than on a schedule, because the rebuild is the event that causes it.”
How would you design a system whose requirements you know will change?
Identify which parts of the requirement are stable and which are volatile, and put the abstraction boundary between them rather than everywhere. Buy flexibility where you can name the change you expect — a second provider, a second tenant, a second region — and commit hard everywhere else. Speculative generality costs every future reader, and the flexibility you did not need is indistinguishable from complexity you cannot remove.
Hands-on Lab
Section titled “Hands-on Lab”This module has no lab of its own — it is the method the others are instances of. The most direct practice is to reverse the process on a system that already exists.
Take async-ai-gateway and recover its constraints from its code:
what must have been true for per-tenant limiting, circuit breaking, and explicit draining each to be
worth building? Then compare against the
Async AI Gateway architecture page, which states them.
The gap between what you inferred and what was intended is the part of design that does not survive
into code — which is the argument for writing it down.
Model a system you already run
Take a service you operate and fill in the Workload above with its real numbers. Compare the
computed concurrent_requests against its actual connection pool and concurrency limits, and
cost_per_day against the invoice. In most systems at least one of those disagrees with the
configuration by more than a factor of two — and the disagreement is a finding, whichever
direction it goes.
Before calling a design done
- Every constraint is a number or a named party, not an adjective.
- The capacity arithmetic has been done, and it agrees with the configured bounds.
- The decision that is hardest to reverse has been named, and deliberated in proportion.
- Each claim has a measurement that would refute it and a threshold for acting.
- There is a quality signal, not only a latency one.
- The dominant cost line is known, and the design lever that moves it is identified.
- What breaks first at ten times the load has an answer.
- What was deliberately left out of v1 is written down, with why.
References
Section titled “References”- Jeff Bezos, 2015 Amazon shareholder letter — the one-way and two-way door framing, which is the most portable idea in this module.
- Martin Kleppmann, Designing Data-Intensive Applications — the reference for the storage and consistency decisions this method routes into.
- Google, Site Reliability Engineering, chapters on SLOs and error budgets — how a design claim becomes an operational commitment.
- John Ousterhout, A Philosophy of Software Design — the clearest treatment of when an abstraction earns its cost, which is the “optimize for change” trade-off above.
- ADR — this handbook’s own decision records, as worked examples of the written output this module argues for.
Revision History
Section titled “Revision History”| Version | Date | Change |
|---|---|---|
| 1.0.0 | 2026-08-09 | Initial publication. |