Skip to content

Module 0: Principal Engineer Mindset

Listen to this page18:30
Read the transcript

1. Why this module has no tech stack

Host: So every other module in this series is going to hand you a technology — asyncio, Kubernetes, RAG, MCP, LangGraph, the whole stack. This one is different, and I want to be upfront about that right at the start: there’s no tool here, no framework you’re going to install. We’re talking about judgment.

Guest: Right, and the reason we’re front-loading this is that judgment is actually the thing separating a strong Senior Engineer from a Principal Engineer — not deeper technology knowledge. Give a Senior a well-scoped ticket and they’ll produce a good implementation, no question. But give a Principal a vague, possibly wrong problem statement, and their first move isn’t to build — it’s to figure out whether that’s even the right problem to solve. Then they produce a decision that a team of strangers can execute without them in the room, with the trade-offs and blast radius already spelled out, and a signal baked in for when it needs revisiting. That’s the muscle this module is training, and honestly every other module in the handbook assumes you already have it.

2. The Leverage Ladder

Host: Okay, so if judgment is the muscle, what’s the actual shape of it? You mentioned a ladder earlier — walk me through that, because I think people hear ‘leverage’ and think it just means ‘do less work yourself.’

Guest: It’s really about where your time produces value, not how much code you write. A Junior is asking ‘how do I build this correctly,’ their lever is their own code. Mid-level asks ‘how do I deliver this reliably’ over a feature. Senior owns a component and asks how it should behave. Staff owns a system and asks how components work together. Principal owns a platform or org and asks the biggest question of all: what’s the simplest system that keeps working as the company grows. Each rung up, you’re trading direct output for leverage through other people’s work — a Principal who spends a week hand-optimizing one function has usually made a leverage mistake, because that same week as a design review could’ve stopped ten teams from making the same error.

Host: But that sounds like it could tip into its own trap — always reaching for the platform-level fix because that’s what a Principal is ‘supposed’ to do.

Guest: Exactly, and that’s the nuance that trips people up. Reaching for leverage is a lens you apply after you’ve confirmed the problem is real and recurring — not a reflex you pull because it feels more senior. We’ve got a whole section later on premature platform-building, where someone builds elegant infrastructure for a problem that had exactly one caller. The ladder tells you where leverage lives; it doesn’t excuse you from first checking whether you’re solving something that actually needs solving.

3. Treating a decision as a system: Type 1 vs Type 2 doors

Host: So let’s zoom out from leverage for a second and talk about the decision itself as a thing you build. You’ve described it before as having an architecture — input, process, output. What does that actually mean in practice?

Guest: It means you stop treating decisions as events that happen in someone’s head and start treating them as a pipeline. The input is the ambiguous problem, the process is however it actually gets evaluated — even if that evaluation is just a hallway conversation — and the output is a decision record with consequences attached. Most dysfunction I’ve seen traces back to that process step being informal or missing entirely, so six months later nobody can reconstruct why the system looks the way it does.

Host: And inside that process step, you’ve said there’s one branch that matters more than all the others — reversible or not. That’s the Bezos door framing, right?

Guest: Right, Type 1 versus Type 2 doors, from his 2015 letter. Type 2 is reversible — you walk back through it — so it should move fast with async written review from a small group; over-processing those is its own leverage mistake, you’re burning expensive synchronous time on a risk a rollback would’ve solved for free. Type 1 is the data model migration, the public API contract, the vendor lock-in — hard or impossible to reverse — and that earns the synchronous review, the wider stakeholder list, the extra week, because being wrong and stuck costs far more than being slow. Almost every bad technical decision I’ve untangled wasn’t wrong on the merits — it was the right rigor applied to the wrong door: a Type 1 call made in a Slack thread, or a Type 2 call stuck in a month of review.

4. The six-question framework and the seven lenses

Host: So once you know which door you’re walking through, how do you actually structure the decision itself? You mentioned a six-question sequence earlier — walk me through it.

Guest: Problem, constraints, alternatives, trade-offs, decision, consequences — and the order is the whole point. Most bad docs jump straight to alternatives before nailing down constraints, so you get a beautiful comparison of options none of which actually fit. Problem statement can’t name a technology — ‘we need Kafka’ isn’t a problem, ‘tenant A’s spikes are timing out tenant B’ is — and constraints have to surface before you propose anything, because a constraint you discover afterward is one you didn’t bother to look for.

Host: And alternatives means more than one option dressed up to look like a choice.

Guest: Right, at least two, including a boring baseline like ‘just add a queue and change nothing else’ — one option is a decision pretending to ask permission. Then trade-offs, the section everyone skips because it feels like arguing against yourself, followed by a plain decision tied to today’s constraints, not elegance or nostalgia for your last job. And consequences has to name the actual signal that would make you revisit it — otherwise it’s not a decision, it’s dogma. Run every option through all seven lenses, too — customer value, simplicity, reliability, scalability, security, cost, maintainability — because if you only ever ask ‘is this elegant to build,’ you’ll pick the option that’s fun for you and painful for whoever’s on call in a year.

5. Writing it down: decision records and ADR-0003

Host: So all seven lenses run through your head, you land on a decision — what actually comes out the other end? Is there a template, or is this more of a mental checklist you carry around?

Guest: It’s a written artifact, an actual decision record, and this handbook practices what it preaches — every architectural call behind the platform lives in slash-adr. Take ADR-0003, it’s short and shows the six-question shape in miniature: the site needs Mermaid diagrams and dark mode, and Mermaid only knows one theme per render, so the problem is build-time SVG versus client-side rendering that can react to a live toggle. Three options get listed, including the one they rejected — build-time rendering — and why: it needs a headless Chromium dependency in every CI build just to pre-render two SVGs per diagram, which is a real cost, not a hand-wave.

Host: And the decision itself is just one paragraph — render client-side, keep the diagram source in a data attribute so it survives re-rendering, watch the theme attribute, re-render on change. Notice what’s missing too — no tour of Mermaid’s API, no history lesson, just enough context to be checkable. Two things make a record like that actually get used instead of ignored — write it before you’re sure, not after you’ve already shipped, because a record justifying a done deal can’t surface the disagreement that would’ve improved it, and state the revisit trigger as a fact, like ‘if a meaningful share of readers have JS disabled,’ not a feeling like ‘if it stops feeling right’ — that vague version is exactly how bad decisions survive for years.

6. Production walkthrough: the tenant-isolation incident

Host: Let’s run an actual incident through this so it stops being abstract. Gateway serves multiple tenants, tenant A has a traffic spike, and suddenly tenant B’s requests start timing out. A strong Senior sees that and immediately reaches for a per-tenant queue in the request handler — what’s wrong with that instinct?

Guest: Nothing’s wrong with it as code, it’s wrong as a decision, because it skips straight to implementation without asking whether this is a one-off or a pattern this gateway will hit again with other tenants at other scales. So walk it through the six questions instead. Problem: noisy-neighbor latency, and notice it was discovered via a support ticket, not monitoring — that’s a separate decision to revisit later. Constraints: the gateway is stateless behind a load balancer with multiple replicas, so whatever you build has to work globally across replicas, not just inside one process.

Host: And that constraint is exactly what kills the naive fix, right? An in-process token bucket per tenant looks like the boring, safe baseline, but it’s per-replica.

Guest: Right, a tenant effectively gets replica-count times their real quota, which defeats the whole point. So the real alternatives are a Redis-backed limiter shared across replicas, or pushing this to a service mesh with per-tenant rate limiting at the infrastructure layer — rejected for now, too big an organizational lift for what the problem currently justifies. Redis wins, but the trade-off has to be explicit: it adds a dependency and a failure mode, so the decision isn’t just ‘use Redis,’ it’s ‘use Redis, and when Redis is down, fail open with a tighter local emergency limit’ — documented on purpose instead of discovered during the next outage. That’s also the actual shape of the isolation logic in the async-ai-gateway lab, so you can go read the same decision as running code.

7. The framework becomes code: async-ai-gateway

Host: So walk me through the lab itself. You’ve got production_app and secure_app sitting side by side in async-ai-gateway — why not just one app with everything turned on?

Guest: Because collapsing them would bury the identity and quota layer inside every reliability example, and you’d never see it on its own. production_app gives you bounded concurrency, retries, deadlines, health-aware fallback — everything except who’s asking. secure_app adds JWT-verified tenant identity and the Redis-backed limiter, and that limiter is the whole point: it’s a single Lua script doing refill-and-consume atomically, because at capacity twenty with two replicas, a naive get-then-set lets both replicas think they have the full quota and you end up serving forty.

Host: That’s a great one to sit with — the atomic Lua script is exactly what stops the double-quota bug. Before we move on, keep two more tensions from that lab in your head, because we’re about to hit them head-on: semaphores bound concurrency while token buckets bound arrival rate, and readiness is not liveness — conflating either pair is how a correct design still takes down a rolling deploy.

8. When the safety net lies to you: the Redis CI story

Host: So let’s talk about a green checkmark that lied. The async-ai-gateway repo had a Redis integration job in CI — service container spun up, tests selected with a pytest dash k redis filter. Sounds thorough. What was actually wrong with it?

Guest: Every test that filter matched was backed by a FakeRedis stub. It returns a canned answer and never opens a socket, so the container just sat there unused. That job was green whether Redis existed or not — it would have passed if the container had never started at all.

Host: Which means it was certifying nothing about the atomicity claim we just spent all that time on. What actually closed the gap?

9. Five ways this goes wrong

Guest: What actually closed the gap was a test designed to break the guarantee, not just exercise it. The mock never could have caught that because it didn’t have a real transaction boundary to violate. That’s the only kind of test that proves atomicity — one that tries to break it.

Host: Okay, so zooming out from that one incident — you’ve been doing this long enough to see the same mistakes recur across teams. If you had to name the five ways smart engineers derail one of these decisions, what are they?

Guest: First, technology-first design docs — the title is a product name like ‘Our Kafka Migration’ before anyone’s written the problem statement. Second, building a generalized platform for a caller count of one, which is the leverage-ladder mistake in its most common clothing. Third, happy-path-only designs with no failure-mode section, so deployment and on-call become someone else’s problem later. Fourth, optimizing a bottleneck you assumed exists instead of one you measured — the fix is a number, not an instinct. And fifth, the mirror image of all that: treating a reversible Type 2 decision like it needs full committee consensus instead of one written review.

10. Trade-offs the framework doesn’t resolve for you

Host: So given all five failure modes, is the answer just ‘move fast on Type 2, move slow on Type 1’? That feels like it could become its own cargo cult.

Guest: It would be, if that’s where it stopped — the real lever is classifying the door correctly before you pick a speed, since a fast wrong Type 2 costs a rollback but a fast wrong Type 1 can become permanent because nobody has the appetite to redo it. There’s a second tension underneath that: let every team pick its own patterns and you get short-term velocity with long-term fragmentation, but centralize everything and you kill the local judgment that caught the tenant-isolation bug. The resolution most Principals land on is centralizing contracts — APIs, event schemas, SLOs — while leaving implementation to whoever’s closest to the problem. And don’t overcorrect into writing a full ADR for every decision either; a low-blast-radius Type 2 might just be three sentences in a searchable Slack thread — what matters is that the reasoning and the revisit trigger exist somewhere, not the document’s length.

11. Measuring the decision process itself: security and performance as applied to judgment

Host: Let’s turn the lens on the process itself, because you keep using words like ‘security’ and ‘performance’ and I don’t think you mean firewalls and latency dashboards. What do those words mean when the thing you’re measuring is judgment?

Guest: Security here is about the integrity of the decision-making process, not an attack surface. The first risk is a single point of failure in judgment — if only one person understands why a decision was made, that’s a bus-factor risk to the whole system, and it’s a big reason decision records exist, so the reasoning survives that person leaving. The second is groupthink: a Principal’s opinion carries outsized weight in a room, which is exactly why the framework demands written alternatives before the meeting — it’s much harder to anchor everyone on ‘my preferred answer’ when two other real options are already sitting on the page. And there’s a third one people forget: in regulated environments, an auditor can ask ‘why does this system work this way’ years later, and a decision record is an answer while a half-remembered Slack thread is not.

Host: Okay, and performance of the process — what would you actually put a number on?

Guest: Three things. Decision latency: time from the problem surfacing to a written decision — if Type 2 calls in your org take three weeks, the process has become the exact bottleneck it was supposed to prevent. Reversal rate: the fraction of decisions revisited within, say, ninety days — near zero means you’re treating Type 2 doors like Type 1, unnecessarily cautious, while very high means someone’s skipping the constraints step. And review overhead, the meeting-hours spent per decision, which is the one number that should trend down over time as an org’s Type 1/Type 2 instincts get sharper.

12. Scaling the framework from 10 to 200+ engineers, and taking it home

Host: Let’s zoom out to org scale, because those metrics you just gave assume a process exists at all. At ten engineers, do you even need any of this, or is a Slack thread and a nod from the team lead sufficient?

Guest: Totally sufficient, and writing it down at that size is often pure overhead because everyone already shares the context. The trouble starts around fifty to a hundred, when two teams independently make conflicting Type 1 calls solving the same cross-cutting problem two incompatible ways, and nobody notices until they collide. That’s the point where a lightweight, searchable ADR log starts earning its keep — not a heavyweight review board, just visibility. Past two hundred, the Principal’s actual job flips: you stop making most of the decisions yourself and start designing the process other engineers use, the template, which categories need a review board — and per the Leverage Ladder, teaching that framework becomes the highest-leverage thing you can do. Watch for the failure mode where the ten-person process just gets frozen in place: it either bottlenecks through one person or becomes theater nobody actually follows before shipping.

Host: That maps onto something else worth naming before we close — this is exactly what interviewers are listening for, isn’t it, even when they’re not saying ‘tell me about your framework.’

Guest: Exactly, they’re listening for the shape, not the outcome — the constraint that ruled out the popular option, the trade-off you knowingly accepted, and what would’ve changed your mind, stated before things went wrong, not as a retrospective excuse. So here’s the homework: pick one real decision from your own work, even a small one, and write it up with the six questions and the ADR template — use ADR-0002, our own pnpm-workspace decision, as your length reference, because it’s genuinely small and written up properly instead of padded to sound important. If you get to the end and can’t state a concrete, checkable revisit trigger, that’s your signal the constraints step got skipped — fix that, and you’ve got the mindset, not just the module.

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

Every module after this one teaches a technology: asyncio, Kubernetes, RAG, MCP, LangGraph. This module teaches none of those, because the thing that actually separates a Principal Engineer from a strong Senior Engineer isn’t technology knowledge — it’s how they frame ambiguous problems, weigh trade-offs under incomplete information, and make decisions that keep working after they’ve moved on to the next problem.

A Senior Engineer, given a well-scoped ticket, produces a good implementation. A Principal Engineer, given a vague and possibly wrong problem statement, first figures out whether it’s the right problem, then produces a decision — with its trade-offs, its blast radius, and the signal that would tell someone it needs revisiting — that a team of strangers can execute without the Principal Engineer in the room. This module is the decision-making framework the rest of the handbook assumes you already have.

The Leverage Ladder. Every engineering level trades direct output for indirect output through other people’s work:

Level Primary lever Central question
Junior Engineer Their own code “How do I build this correctly?”
Mid-Level Engineer Their own features “How do I deliver this reliably?”
Senior Engineer A component “How should this service behave?”
Staff Engineer A system “How should these components work together?”
Principal Engineer A platform or org “What’s the simplest system that keeps working as the company grows?”

The ladder isn’t about title politics — it’s about where your time produces the most value. A Principal Engineer who spends a week hand-optimizing one function has, in most orgs, made a leverage mistake: that week would pay off more as a design review that prevents ten teams from making the same mistake, or a library that removes the need for the optimization everywhere at once. The skill this module builds is recognizing which lever you’re actually pulling, and choosing the highest-leverage one available — not writing the most code, or making the most decisions, but making the decisions that compound.

Leverage cuts both ways

The same instinct that makes you effective — reaching for the platform-level fix instead of the local patch — is also the failure mode this module spends a whole section on later: premature platform-building for a problem that had exactly one caller. Leverage is a lens you apply after you’ve confirmed the problem is real and recurring, not a default you reach for because it feels more “Principal.”

Treat a technical decision itself as a system with an architecture: an input (an ambiguous problem), a process (how it gets evaluated), and an output (a decision record with defined consequences). Most engineering-culture dysfunction traces back to a missing or informal version of this pipeline — decisions made in hallway conversations, never written down, with no way for someone six months later to know why the system looks the way it does.

The two branches after “reversible?” are the single highest-leverage distinction in this architecture. It comes from Jeff Bezos’s 2015 shareholder letter, and it should shape how much process a decision gets:

  • Type 2 decisions (reversible, one-way doors you can walk back through) should move fast, with async written review from a small number of people. Over-processing a Type 2 decision is a leverage mistake in the other direction — you’ve spent expensive synchronous time protecting against a risk that a rollback would have solved cheaper.
  • Type 1 decisions (hard or impossible to reverse — a data model migration, a public API contract, a vendor lock-in) deserve the synchronous design review, the wider stakeholder list, and the extra week. The cost of being slow is real but bounded; the cost of being wrong and stuck is not.

Most bad technical decisions aren’t wrong on the merits — they’re the right amount of rigor applied to the wrong door type: a Type 1 decision made in a Slack thread, or a Type 2 decision stuck in a four-week review cycle.

The pipeline above needs a concrete framework to fill in the “process” box. Use these six questions, in order, for any non-trivial decision — the order matters, because skipping ahead to “alternatives” before “constraints” is the most common way design docs go wrong:

  1. Problem. State the user, business, and engineering problem without naming a technology yet. “We need Kafka” is not a problem statement; “tenant A’s traffic spikes are causing tenant B’s requests to time out” is.
  2. Constraints. Latency, scale, security, compliance, cost, deadline, team skill set, and what’s already load-bearing in production. Constraints you discover after proposing a solution are constraints you didn’t do the work to find.
  3. Alternatives. At least two real options, one of which is a deliberately boring baseline (“do nothing new, just add a queue in front of the existing service”). A design doc with one option isn’t a design doc — it’s a decision looking for a rubber stamp.
  4. Trade-offs. What each option improves, what it costs, and what it makes harder to change later. This is the section people skip because it feels like arguing against your own proposal — it’s also the section that makes the eventual decision defensible in six months.
  5. Decision. State it plainly, tied to the constraints that made it the right fit now — not “the most elegant architecture” or “what the team used at their last job.”
  6. Consequences. Risks accepted, follow-up work, the rollback path, and — critically — the concrete signal that would justify revisiting the decision. Without this, decisions calcify into dogma because nobody knows what evidence would change their mind.

Research Note

The Type 1/Type 2 framing is worth reading in the original — it’s two paragraphs, and it’s the single most load-bearing idea in this module.

Source: Jeff Bezos, 2015 Amazon shareholder letter

Evaluate every option against the same seven lenses, every time, so you don’t accidentally optimize one dimension (usually “elegance” or “what’s fun to build”) at the expense of the others:

The seven lenses

  • Customer value — does this solve the right problem for the right user? - Simplicity — is this the least complex design that satisfies the actual requirements? - Reliability — how does it fail, recover, and degrade? - Scalability — what breaks at 10x traffic, data, tenants, or team size? - Security — are identity, authorization, data protection, and audit built in, not bolted on? - Cost — infrastructure, model/API, operational, and staffing cost, not just the cloud bill - Maintainability — can a team that didn’t build this operate and evolve it?

“Implementation” for a decision-making framework means the artifact you actually produce, not code. The artifact is a written decision record — and this handbook dogfoods exactly the format it’s teaching. Every architectural decision behind this platform itself lives in /adr/, using the same six-question shape as a real, non-hypothetical example:

Read a real ADR as a worked example

ADR-0003 is a good one to start with — it’s short, and it shows the shape in miniature: Context (why this came up), Problem (the specific question), Options (three of them, including the one that got rejected and why), Decision (one paragraph), and Consequences (what it costs, including “if this assumption stops holding, revisit it”). Notice what’s absent: no technology tour, no unrelated background — just enough context for the decision to be checkable.

Two implementation details make the difference between a decision record that gets used and one that gets ignored:

  • Write it before you’re sure, not after. A decision record written to justify a decision that’s already shipped is documentation theater — it can’t surface the disagreement that would have improved it. Circulate the draft before you start building.
  • State the revisit trigger as a fact, not a feeling. “Revisit if p99 latency exceeds 400ms” is checkable. “Revisit if this doesn’t feel right anymore” is not — and it’s how technically bankrupt decisions survive for years past their expiration date.

A tenant-isolation incident, walked through the framework above:

A gateway serving multiple tenants starts seeing tenant B’s requests time out during tenant A’s traffic spikes — the exact scenario used as the problem statement above. A Senior Engineer’s first instinct is usually the local fix: add a per-tenant queue in the request handler. That’s not wrong, but it’s an implementation, not a decision — it hasn’t asked whether “per-tenant isolation” is a problem this gateway will face again (across which other tenants, at what scale) or a one-off.

Run it through the framework:

  1. Problem: noisy-neighbor traffic causes cross-tenant latency, discovered via a support ticket, not via monitoring — itself a signal worth a separate decision later.
  2. Constraints: existing gateway is stateless behind a load balancer with multiple replicas; any per-tenant limiter has to work correctly across replicas, not just within one process.
  3. Alternatives considered: (a) in-process token bucket per tenant — the boring baseline, but wrong here because it’s per-replica, not global, so a tenant gets replica_count × their real quota; (b) a Redis-backed distributed limiter shared across replicas; (c) push isolation to the infrastructure layer (a service mesh with per-tenant rate limiting) — rejected for now because it’s a bigger organizational lift than the problem currently justifies.
  4. Trade-offs: (b) adds a Redis dependency and a failure mode — what happens to tenant isolation if Redis itself is down — that (a) doesn’t have.
  5. Decision: ship (b), and make the Redis-unavailable behavior an explicit, documented choice (fail-open with a tighter local emergency limit) rather than an accident.
  6. Consequences: on-call now owns a new dependency; revisit if Redis latency becomes a measurable fraction of request latency, or if a tenant tier needs stronger isolation than a shared limiter can give.

This is the actual shape of the decision implemented in labs/async-ai-gateway — see its distributed rate limiting and the fail-open/fail-closed trade-off called out in that lab’s documentation. The mindset framework and the production code are the same decision, described twice, once as reasoning and once as implementation.

Technology-first design

Picking Kafka, an agent framework, or a vector database before the problem statement exists. You can tell this has happened when the design doc’s title is a product name instead of a problem.

Premature platform-building

Building a generalized, multi-tenant, config-driven version of something that has exactly one caller today. This is the leverage-ladder mistake from the Mental Model section, in its most common concrete form.

Ignoring operations

Designing the happy path and treating deployment, monitoring, rollback, and on-call as an exercise for whoever ships it. If the design doc has no failure-mode section, this has already happened.

Unmeasured optimization

Optimizing a bottleneck you assumed exists instead of one you measured. Ties directly to the Performance section below: the fix for this failure mode is a number, not an instinct.

Consensus paralysis

Treating “everyone agrees” as the bar for a Type 2 decision that should have moved on one written review. This is the mirror image of technology-first design — instead of too little process, it’s too much, applied to the wrong door type.

Speed vs. certainty

A Type 2 decision made fast and wrong costs a rollback. A Type 1 decision made fast and wrong costs a migration, or worse, becomes permanent because nobody has the appetite to redo it. The lever isn’t “always move fast” or “always be careful” — it’s correctly classifying which door type you’re behind before choosing a speed.

Autonomy vs. alignment

Letting every team choose its own patterns maximizes short-term velocity and long-term fragmentation. Centralizing every choice maximizes consistency and kills the local judgment that caught the tenant-isolation bug in the Production Example. Principal Engineers usually resolve this by centralizing contracts (APIs, event schemas, SLOs) while leaving implementation to the team closest to the problem.

Documentation overhead vs. decision velocity

Writing a full ADR for every decision, including ones nobody will ever question, is its own failure mode. The six-question framework scales down: a Type 2 decision with low blast radius might get three sentences in a Slack thread with a permalink saved somewhere searchable — what matters is that the reasoning and the revisit trigger exist somewhere, not the document’s length.

Security here means the integrity of the decision-making process itself, not a system’s attack surface:

  • Single point of failure in judgment. A decision that only one person understands is a bus- factor risk to the system, not just the team. Decision records exist partly so a decision survives its author leaving.
  • Groupthink and unchecked authority. A Principal Engineer’s opinion carries outsized weight in a room; that’s exactly why the framework asks for written alternatives before the meeting — it’s much harder to anchor a room on “my preferred answer” if two other real options are already on the page.
  • Audit trail as a compliance artifact. In regulated environments, “why does this system work this way” is a question an auditor can ask years later. A decision record is the answer; a half-remembered Slack thread is not.

Applied to the decision process itself, not a running system, “performance” means:

  • Decision latency — time from problem surfacing to a written decision. A useful thing to actually track: if Type 2 decisions in your org take three weeks, the process itself has become the bottleneck the framework was supposed to prevent.
  • Reversal rate — the fraction of decisions revisited within, say, 90 days. Near zero suggests excessive caution (Type 2 decisions being treated like Type 1); very high suggests the “constraints” step is being skipped.
  • Review overhead — synchronous meeting-hours spent per decision. This is the number that should go down over time as an org’s Type 1/Type 2 instincts improve.

What works for one team of eight breaks past a few hundred engineers, in a predictable order:

  • At 10 engineers: verbal agreement and a shared Slack channel are enough process. Writing things down feels like overhead because everyone already has the context.
  • At 50–100 engineers: teams start making conflicting Type 1 decisions independently — two services solve the same cross-cutting problem two incompatible ways — because no one team has visibility into what the others are deciding. This is where a lightweight, searchable ADR log (not a heavyweight approval process) starts paying for itself.
  • At 200+ engineers: the Principal Engineer’s job shifts from making decisions to designing the decision-making process other engineers use — writing the ADR template, defining which decision categories need a design review board, and, per the Leverage Ladder, teaching the framework itself becomes the highest-leverage lever available.

The scaling failure to watch for: a process that worked at 10 people, kept unchanged at 200, becomes either a bottleneck (every decision routes through one person) or theater (the process exists on paper but nobody actually uses it before shipping).

Tell me about a technically unpopular decision you made.

Interviewers are listening for the six-question shape, not the outcome. A strong answer states the problem, the constraints that ruled out the popular option, and — critically — what consequence you accepted and what would have changed your mind. A weak answer is a war story with no mention of what you’d have needed to see to decide differently.

Model answer shape

“We were migrating [system] and the team wanted [popular option] because [reason]. I proposed [less popular option] because [constraint the popular option didn’t satisfy]. The trade-off was [cost we accepted] in exchange for [benefit]. We set [specific metric] as the signal to revisit, and [it did / didn’t] come up.” Notice this shape works whether the decision was later validated or not — the framework, not the outcome, is what’s being evaluated.

How do you decide whether to write an RFC/ADR versus just building it?

Map directly to Type 1 vs. Type 2 from the Architecture section: reversibility and blast radius, not personal preference or how interesting the problem is.

Describe a decision you got wrong. What did you learn?

This tests whether the candidate treats “consequences” as a real section or an afterthought. Listen for whether they had defined a revisit trigger before things went wrong — if the answer is “we just noticed it wasn’t working,” the framework in this module wasn’t actually being used at the time.

How do you influence a team without direct authority over them?

Ties to the Leverage Ladder: authority-free influence comes from writing the option comparison other teams reuse, not from being the person in the room with the most seniority. A concrete example beats a values statement here.

Write ADR-000X for a scenario you've actually seen

Pick a real decision from your own work — even a small one — and write it up using this module’s six-question framework and the ADR template convention in docs/CONTENT_GUIDE.md. Use ADR-0002 as a length/depth reference — it’s a genuinely small decision, written up properly, not padded to look more important than it is. The exercise isn’t complete until you’ve written a concrete, checkable revisit trigger; if you can’t state one, that’s a sign the “constraints” step was skipped.

No code lab for this module

Unlike every module after this one, Module 0 has no companion service in labs/ — the artifact this module produces is a decision record, not code. Module 1 picks back up with a real, running system.

  • Jeff Bezos, 2015 Amazon shareholder letter — the original Type 1/Type 2 decision framing.
  • Will Larson, Staff Engineer: Leadership Beyond the Management Track — the leverage-ladder framing of Staff/Principal work, and concrete advice on writing technical documents that get read.
  • Tanya Reilly, The Staff Engineer’s Path — especially its treatment of “big-picture thinking” and execution without direct authority.
  • Martin Fowler, “Technical Debt Quadrant” — a useful companion lens for the Trade-offs section.
  • Michael Nygard, “Documenting Architecture Decisions” — the original ADR proposal this platform’s own /adr/ section is built on.
Version Date Change
1.0.0 2026-08-05 Initial publication, expanded from the retired static-prototype draft.