Module 14: Leadership
Read the transcript
1. The constraint changes at Principal scope
Host: So we’re starting Module 14 on leadership, and I want to push back on the word a little, because usually when people say leadership they mean something soft bolted onto the engineering. That’s not what this module is arguing.
Guest: Right, the claim is narrower and more mechanical than that. At Principal scope your constraint stops being what you personally can build and becomes what you can cause to be built correctly by people who don’t report to you. The output of that work is still concrete — it’s decisions, and the mechanisms that make those decisions actually stick after you leave the room.
Host: And you’re saying AI platforms are one example of where that shows up.
Guest: Exactly, two things stack up there. Research and product genuinely disagree on what ‘done’ means, so someone has to force that decision explicitly instead of letting it stay implicit. And these systems degrade silently, so unless someone insisted on measurement before anyone felt pain, you find out you were wrong months too late.
2. Decisions and mechanisms, not opinions and reminders
Host: So if the job is forcing those decisions and insisting on measurement, what’s actually different about the artifact you produce versus what you produced as a senior engineer? It sounds like you’re still just… having opinions, but louder.
Guest: The difference is whether it has to be re-litigated. An opinion gets re-argued every time someone new joins the team, because there’s nothing to point to. A recorded decision, with the options that lost and what would change the answer, gets read once and the question stays settled. Same with reminders versus mechanisms — a reminder decays the week you stop saying it in standup, but a mechanism that fails the build doesn’t care whether you’re in the room.
Host: Okay, so the actual output is the decision record and the mechanism, not the correct take. What does that imply for how you spend your time day to day?
Guest: Three things follow. First, route by reversibility, not importance — most calls are cheap to undo and should stay with whoever’s closest to the code, or you become the bottleneck. Second, write it down if you want it to outlive the meeting, because writing forces the constraints into the open for people who weren’t there, including future-you. Third, a norm nobody enforces is just a preference — ‘we document decisions’ means nothing until there’s a CI check that fails without one.
3. Routing by reversibility, not importance
Host: So walk me through the actual mechanics of that first move — route by reversibility. What does ‘cheap to undo’ look like in practice versus the stuff that actually deserves a Principal’s attention?
Guest: Ask one question of every decision: if this is wrong, what does fixing it cost? A retry policy is a config change, so it stays with whoever’s closest to the code — pulling that into review just makes you the bottleneck while feeling productive. A queue technology choice costs a week, which is worth a conversation. But a tenant data boundary, if it’s wrong, costs a migration, an audit, and a disclosure — that’s the one that earns your time and a written record.
Host: That’s the exact same ladder from the systems design module, just pointed at org decisions instead of one architecture.
Guest: Right, it’s literally the same test — Module 13 ranks choices inside a single design by that cost-to-undo question, and here you’re ranking decisions across a whole team by it. The insight doesn’t change scale, it just changes what’s being routed: instead of where you spend design effort, it’s where you spend organizational attention, and most of it should never reach you at all.
4. Disagreement that converges
Host: So say I’ve routed correctly and something does land in front of me — two senior engineers who genuinely disagree about a design. Where do you even start unpacking that?
Guest: You figure out which of three things it actually is, because almost every technical argument is one of them. Constraint disagreement, conclusion disagreement, or preference — and they don’t get solved the same way, so misdiagnosing which one you’re in is why arguments run for weeks.
Guest: Constraint disagreement is when one person’s building for 200 requests a second and the other’s building for 20 — that’s not an architecture debate at all, it just looks like one until someone writes the number down. Conclusion disagreement is the legitimate kind, same constraints, different read on the evidence, and you resolve that by naming what evidence would settle it and what it costs — a two-day spike beats a two-week argument. Preference is same constraints, same evidence, just different taste, and that’s a two-way door: whoever owns the code decides, full stop.
Host: And I’d guess the expensive mistake is spending senior attention on the last one, arguing preference like it’s a conclusion disagreement.
Guest: Exactly — that’s the most common way seniority gets wasted, litigating taste as if it were substance. It shows up in interviews too: the answer that lands isn’t ‘I made the right call,’ it’s naming which of the three you were in, what evidence you bought, and how long the gap was between the evidence arriving and the decision changing. Everyone’s been wrong — the interesting number is how fast you noticed.
5. Mechanisms over vigilance, made concrete
Host: So let’s make that concrete, because ‘automate it instead of reminding people’ is easy to nod at and hard to picture. What does that actually look like on your team?
Guest: The clearest one is our ADR check. Every architecture decision record has to carry the same seven sections — Status, Context, Problem, Options, Decision, Consequences, References — and CI fails the build if one’s missing. No reviewer has to remember to ask ‘did you consider alternatives,’ the pipeline just refuses to merge until the section exists.
Host: Why those two fields, Options and Consequences, in particular — what breaks if you drop them?
Guest: Options forces the rejected alternatives onto the page, which is the only way to later tell a considered decision from a default nobody thought about. Consequences forces the cost to be written by the person who’s most honest about it, which is right now, not six months later when it’s inconvenient. And the check never grades whether the content is good — it just enforces the shape, because a mechanism that tried to judge quality would be unusable. The real test is that it’s failed on my own work — it rejected a page of mine mid-session for a renamed heading, and that’s the only evidence it’s actually doing something.
6. The number that ends the argument
Host: You mentioned mechanisms replacing arguments — I want the clean example, because ‘ship faster versus be more reliable’ feels like it should be a permanent argument. What’s the number that ends it?
Guest: Error budget. Both sides agree in advance on how much unreliability is acceptable over a window. Budget remaining means ship, take the risk, move fast. Budget exhausted means stop, the next unit of work is stability, full stop. Nobody has to win the argument about whether reliability matters more than velocity, because the number already encodes the trade both sides accepted before there was a specific incident to be emotional about.
Host: And the version of that tension specific to working on an AI platform — where does the same kind of number need to exist?
Guest: Research and product don’t even agree on what ‘done’ means — a model that gains two benchmark points is a result to research and a migration to product, who now has to requalify everything built on top of it. Platform sits in the middle, and the job is making explicit what’s guaranteed to stay stable, what isn’t, and who eats the cost when a model shifts under a product that assumed it wouldn’t. Without that written down, every model update becomes the ship-versus-reliability argument again, except nobody agreed on a number this time.
7. The reversal: verify, decide, then record why
Host: So walk me through the moment this handbook’s own direction got reversed. That’s a good story to end this argument on.
Guest: We’d built a lot of content on a static-HTML architecture while a separate branch carried a documentation platform. An outside review came in and said the static site was the retired one, don’t merge it forward. Taken at face value that costs you weeks of work; dismissed reflexively, you keep building on a wrong foundation.
Host: So which was it — did the review hold up?
Guest: Partly. We checked every claim against the repo instead of against the argument, and two of them were just wrong — the lab migration wasn’t mechanical since only one lab had production artifacts, and the actual hazard was that main was still serving the retired site, not the branch itself. The rest of the review held and we changed direction, but the part that mattered was writing down which claims didn’t survive, not just which ones did — because a decision that only records what was accepted reads six months later like there was never any doubt, and that’s what makes the next person afraid to challenge it.
8. Where leadership fails quietly
Host: So that’s one way a decision quietly rots. What’s the version where the whole leadership function rots, not just a single call?
Guest: It starts with routing everything through yourself, and it feels like diligence while it’s happening. The tell is a queue of small reversible decisions with your name attached to all of them — you’re not the bottleneck because you’re slow, you’re the bottleneck because nobody else was told they’re allowed to own that class of decision. Say it once, publicly, and the queue drains.
Host: And while you’re clearing that queue, attention just goes to whatever’s loudest — this week’s incident, the team that pings you the most.
Guest: Right, and reversibility doesn’t correlate with urgency. Same pattern shows up as unfalsifiable strong opinions nobody can push back on, norms like ‘we always write ADRs’ that hold for about a quarter after the person stops repeating it, and in AI systems specifically — model quality has no pager, no owner, so it just degrades for weeks while everyone assumes someone’s watching it.
9. Deciding early vs. waiting, and how much to standardize
Host: So once you’ve sorted a decision into two-way or one-way door, how do you actually stop ‘let’s gather more evidence’ from becoming a permanent state on the one-way ones?
Guest: You timebox it explicitly, because otherwise more evidence isn’t a step toward a decision, it’s a substitute for one. The cost of waiting is real but invisible, nobody logs the weeks a team sat blocked, so it never shows up on the ledger the way a wrong call does. Bounding it forces the trade-off into the open instead of letting caution masquerade as diligence.
Host: That same instinct — pick a few things to lock down and let the rest breathe — is that also how you’d think about standardizing across teams, like a shared stack versus letting each team choose its own?
Guest: Exactly the same shape. Full autonomy means every team optimizes for its own problem and you end up operating ten different systems; full standardization means one runbook but a straitjacket for anything unusual. So you name the small set that has to be consistent — identity, observability, deployment — and everything else, a team’s storage choice, their internal service boundaries, is theirs to vary.
10. Ownership as a design decision: security and approval authority
Host: You said the small set that has to be consistent includes identity and observability — that sounds like it lands directly on security. Where does ownership as a leadership decision actually show up there?
Guest: Right at the point of who gets to say yes. Who can grant an agent tool access, widen a data boundary, or ship a model that touches customer data — those need named owners before an incident, not during one. There’s a technical half to this, the policy-gated tool execution architecture, which enforces capability scope, argument validation, per-tool quotas, human approval on high-risk actions, and a complete audit trail tied to a verified identity. But none of that matters if the organizational half is missing — if nobody can say, definitively, I am the person who approves this. And the second failure mode is timing: a security review that shows up after the data boundary is already set can only document risk, it can’t remove it. Pulling threat modelling into design review instead of pre-launch is a process change, and it’s the kind only a senior engineer has the standing to push through.
Host: So if someone’s sitting in that design review and the conversation is going nowhere — how do you make it concrete instead of a vague back-and-forth about whether something feels secure?
Guest: You turn it into a scoping decision: does this thing need access to the whole database or just one table, does this credential need to reach production or just staging. Make the secure path the easy path and you stop depending on someone remembering to look.
11. Time-to-decision and mechanisms that scale past you
Host: That scoping habit is a mechanism too, and it makes me want to ask about something people almost never measure: how long a decision sits before it’s made. Is that actually a number worth tracking, or is it just a vibe people complain about at retros?
Guest: It’s a real number and it’s usually a worse bottleneck than anyone admits. A design decision stuck for three weeks is three weeks of a team’s throughput gone, and it never shows up on a dashboard because nothing labeled ‘waiting’ ever does. Same with review latency — a pull request sitting for a day doesn’t cost a day, it costs more, because both sides have to rebuild context before they can even resume. That’s why written, async proposals usually beat a meeting for anything that needs real consideration: they scale to people who weren’t in the room, they produce the record for free, and they surface objections from the person who’d never interrupt out loud. And if you want a gut check on meetings themselves, just count them — eight people for an hour is a working day, and most recurring meetings have never been weighed against that bar.
Host: So fast decisions and fast review are leadership behaviors, not just nice-to-haves. Where does that connect to scaling past yourself?
Guest: Directly — mechanisms scale, your attention doesn’t, and every check that runs without you is capacity you get back. That’s why you delegate decision rights, not tasks: ‘own this component, including what goes in it’ keeps paying off, while ‘do this task’ just needs you again next week. Written artifacts are the same idea stretched across time instead of headcount — you’re writing for the engineer who joins in a year with none of your context and no one left to ask, and the measure of your leadership is what keeps working after you’ve moved on to the next problem.
12. Closing synthesis: what this looks like in practice
Host: So if someone wanted the whole module compressed into one interview answer, what does that sound like?
Guest: That’s really the same ground we already covered — the constraint-conclusion-preference split, encoding the decision so it fails without a person watching, and naming the gap when you’re wrong instead of just admitting you were wrong. The interview-answer version is just that mechanism, said out loud in under a minute, not a new idea.
Host: And there’s no lab bolted onto this one to go practice that on code.
Guest: Right, deliberately — the subject here is the decisions around the code, not the code itself, so the practice is the ADR section, this repository actually doing what the module argues, superseded decisions included. Go read one, then go write one from something your own team decided last year; if you can’t reconstruct what was rejected and why, that decision was never really auditable. That’s the whole arc, and it’s on you now.
Not covered
The planner wanted these and found nothing in the source to support them:
- A deep dive into any specific AI platform incident with named companies or products beyond what’s described generically
- Concrete numeric error-budget policy details beyond the ship/stabilize framing given
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Executive Summary
Section titled “Executive Summary”At Principal scope the constraint stops being what you can build and becomes what you can cause to be built correctly by people who do not report to you. That is not a soft skill wrapped around the engineering — it is engineering work with a different output: decisions, and the mechanisms that make decisions stick.
This module covers that work concretely. How to route a decision by how hard it is to reverse, how to disagree in a way that converges, and how to encode a conclusion so it survives without you repeating it. AI platforms sharpen all of this, because they sit between research and product with genuinely different definitions of “done”, and because a system that degrades silently needs someone to have insisted on measurement before anyone felt the pain.
Mental Model
Section titled “Mental Model”Your output is decisions and mechanisms, not opinions and reminders.
An opinion has to be re-litigated every time a new person joins. A recorded decision — with the options that lost and what would change the answer — is read once and settles the question. A reminder decays the week you stop giving it. A mechanism that fails the build does not.
Three consequences:
- Route by reversibility, not by importance. Most decisions are cheap to undo and belong with whoever is closest to the code. A few are expensive for years. Treating everything as significant makes you the bottleneck; treating nothing as significant is how a data boundary gets decided in a standup.
- Written beats spoken, for anything you want to outlive the meeting. Not ceremony — writing forces the constraints into the open, and it is the only form that scales to people who were not in the room, including your future self.
- A norm nobody enforces is a preference. “We document decisions” is a preference. A CI check that fails when an ADR is missing its Consequences section is a mechanism. Convert the ones that matter; leave the rest as preferences and stop mentioning them.
Influence is a function of being right in a way others can check
Authority makes people comply; a checkable argument makes them agree, which is the only version that holds when you are not in the room. That means showing your constraints, naming what would change your mind, and being visibly willing to update. Engineers who are right and unfalsifiable get routed around within two quarters — not because the org dislikes rigor, but because nobody can tell their strong claims from their weak ones.
Architecture
Section titled “Architecture”How a decision should move, and where most of them should stop early.
flowchart TD
SIGNAL[A problem surfaces<br/>incident, cost spike, blocked team] --> FRAME[Frame it as a decision<br/>what question needs answering]
FRAME --> REVERSIBLE{Reversible in<br/>one deploy?}
REVERSIBLE -->|Yes| OWNER[The team closest to it decides<br/>no review, no meeting]
OWNER --> SHIP[Ship and measure]
REVERSIBLE -->|No| WRITE[Write it down<br/>context, options, cost of each]
WRITE --> CIRCULATE[Circulate for written objection<br/>async, timeboxed]
CIRCULATE --> DISAGREE{Objection<br/>on constraints<br/>or conclusions?}
DISAGREE -->|Constraints| RESOLVE[Resolve the constraint first<br/>most arguments end here]
DISAGREE -->|Conclusions| EVIDENCE[Name the evidence that<br/>would settle it, and its cost]
DISAGREE -->|Neither| DECIDE
RESOLVE --> WRITE
EVIDENCE --> SPIKE[Timeboxed spike]
SPIKE --> DECIDE[Decide and record<br/>including what was rejected]
DECIDE --> ENFORCE[Encode it in a mechanism<br/>linter, test, review gate]
ENFORCE --> SHIP
SHIP --> REVISIT{Evidence contradicts<br/>the record?}
REVISIT -->|Yes| WRITE
REVISIT -->|No| HOLD([Decision holds])The left branch is the one people under-use. If a decision can be undone in a deploy, the correct process is no process — and a Principal engineer who pulls those into review is destroying throughput while feeling productive. The right branch is where the writing, the disagreement, and the mechanism belong.
Deep Dive
Section titled “Deep Dive”Reversibility as the routing key. Ask what fixing it costs if it is wrong. A retry policy is a config change. A queue technology is a week. A tenant data boundary is a migration, an audit, and a disclosure. Rank by that, then spend attention in the same order — Module 13 applies the same test inside a single design.
Disagreement that converges. Almost every technical argument is one of three things, and naming which one collapses most of them:
- A constraint disagreement. You think it must serve 200 requests per second; they think 20. You are not arguing about architecture at all. Resolve this first — a startling share of design arguments end the moment someone writes the requirement down.
- A conclusion disagreement from shared constraints. Legitimate, and resolvable by naming the evidence that would settle it plus what that evidence costs. A two-day spike beats a two-week argument, and beats a two-year wrong decision.
- A preference. Same constraints, same evidence, different taste. This is a two-way door: let whoever owns the code decide and move on. Spending senior attention here is the most common way it gets wasted.
Mechanisms over vigilance. Anything you find yourself saying twice is a candidate for automation. The question is not “how do I get people to remember this” but “what would make forgetting it impossible”. A linter, a required section, a test that fails, a default that is already correct. This is the highest-leverage move available at Principal scope and the one most often skipped in favor of a document nobody re-reads.
The AI-platform-specific tension. Research and product measure “done” differently — a model that improves a benchmark by two points is a result to one and a migration to the other. Platform work sits between them, and the recurring leadership job is making that boundary explicit: what the platform guarantees, what it does not, and who absorbs the cost when a model changes underneath a product that assumed it was stable.
Research Note
The clearest worked example of a mechanism replacing an argument. “Ship faster” versus “be more reliable” is unwinnable as a debate and trivial as arithmetic once both sides agree on a number: budget remaining means ship, budget exhausted means stabilize. The lesson generalizes well beyond reliability — find the number both sides would accept, and the recurring argument disappears.
Source: Google, Site Reliability Engineering — the error budget chapter
Implementation
Section titled “Implementation”The mechanism this repository uses to keep decisions recorded, rather than reminding people to record them. Every ADR must carry the same seven sections, and CI fails if one is missing:
A norm converted into a mechanism, from packages/shared and scripts/
/** Sections every ADR must contain. */export const REQUIRED_ADR_SECTIONS = [ "Status", "Context", "Problem", "Options", "Decision", "Consequences", "References",] as const;
// scripts/lint-content-structure.tsconst CHECKED_SECTIONS: CheckedSection[] = [ { directory: "learn/modules", required: REQUIRED_MODULE_SECTIONS }, { directory: "architecture/systems", required: REQUIRED_ARCHITECTURE_SECTIONS }, { directory: "adr/decisions", required: REQUIRED_ADR_SECTIONS }, { directory: "reference/lookups", required: REQUIRED_REFERENCE_SECTIONS },];
for (const file of files) { const missing = findMissingSections(extractHeadings(stripFrontmatter(source)), required); if (missing.length > 0) { failures.push(`${relativePath}\n missing sections: ${missing.join(", ")}`); }}The interesting entries are Options and Consequences. Requiring Options forces the rejected alternatives to be written down, which is what makes a decision auditable later — without it you cannot tell a considered choice from a default. Requiring Consequences forces the cost to be stated by the person best placed to know it, at the moment they are most honest about it.
Note what this does not do: it never checks whether the content is any good. A mechanism that tried would be unusable. It enforces the shape, and the shape is what makes a bad decision visible during review.
Engineering Note
The test of a mechanism is whether it fails loudly on your own work. This one has — it rejected a page of mine mid-session for a renamed heading, which is the only evidence that it is doing anything. A check that has never failed is indistinguishable from a check that cannot fail. Module 12 takes the same argument into CI, where a pipeline that skipped its tests and one that passed them emit an identical green signal.
Production Example
Section titled “Production Example”This handbook’s own direction was reversed by a review, and the sequence is a usable template.
The project had built substantial content on a static-HTML architecture while a separate branch carried a documentation platform. An outside review argued the static site was the retired one and should not be merged forward. That claim, accepted uncritically, would have discarded weeks of work; rejected reflexively, it would have compounded a wrong foundation.
What actually happened: each claim was checked against the repository rather than the argument. Two
of the review’s assertions turned out to be wrong — the lab migration was not mechanical, because
only one lab had production artifacts, and the real hazard was not the branch but that main still
served the retired site. The rest held, and the direction changed.
The transferable part is the order. Verify the claims, then decide, then record why — including which parts of the incoming argument did not survive. A decision that only records what was accepted reads, six months later, as though there was never any doubt, which is exactly the impression that makes the next person afraid to challenge it.
The reversal is the cheap part
Changing direction cost less than the weeks of content built before anyone asked whether the foundation was right. The leadership failure was not the late reversal — it was that no one had written down which architecture was canonical, so both could look correct at the same time.
Failure Modes
Section titled “Failure Modes”Becoming the bottleneck
Routing every decision through yourself feels like diligence and reads as throughput loss to everyone waiting. The tell is a queue of small reversible decisions with your name on it. The fix is explicit: name the classes of decision teams own outright, and say so once, publicly.
Deciding by proximity instead of by cost
Attention drifts to whatever is loudest — the incident this week, the team that asks most often. Meanwhile the one-way door gets walked through quietly by someone who did not know it was one. Reversibility is the sort key, and it does not correlate with urgency.
Being right in a way nobody can check
Strong conclusions with unstated constraints and no falsifying evidence. People cannot evaluate the claim, so they either defer or route around you, and both are bad. State the constraint, state what would change your mind.
Norms without mechanisms
“We always write ADRs” holds for about a quarter after the person saying it stops saying it. If a norm matters, it needs to fail a build; if it does not matter enough to automate, it probably does not matter enough to repeat.
Hoarding the interesting work
The senior engineer who takes every hard problem produces a team that cannot handle hard problems and a single point of failure with a calendar. The work to delegate is specifically the work you would enjoy — that is the signal that it is growth for someone else.
Letting quality regressions have no owner
In AI platforms this is the one that bites. Latency has an owner and a pager. Model quality often has neither, so it degrades for weeks with everyone assuming someone else is watching. Naming an owner and a threshold is a leadership act before it is a technical one.
Trade-offs
Section titled “Trade-offs”Decide now, or gather more evidence
Deciding early keeps people unblocked and risks being wrong in a way that compounds. Waiting buys accuracy and costs momentum, and the cost is often invisible because nobody logs time spent blocked. Route by reversibility: two-way doors decide now, one-way doors buy the evidence — and bound the gathering with a timebox, or “more evidence” becomes the decision.
Consistency across teams, or team autonomy
A shared stack means one set of runbooks and engineers who can move between services; autonomy means each team fits its own problem and the operational surface multiplies. Set the small number of things that must be consistent — usually identity, observability, and deployment — and let everything else vary.
Fix it yourself, or coach someone through it
Fixing it is faster today and buys nothing tomorrow. Coaching costs your time and theirs, and is the only version that scales past one person. The line worth holding: fix it yourself only when the deadline is real and the lesson is small.
Enforce with a mechanism, or trust the norm
Mechanisms are reliable and rigid — a linter that is wrong 5% of the time gets disabled entirely, taking the other 95% with it. Norms flex but decay. Automate the invariants where a violation is unambiguous and costly, and leave judgement calls to review.
Security
Section titled “Security”- Approval authority is a design decision. Who can grant an agent tool access, widen a data boundary, or ship a model that touches customer data — these need named owners before an incident, not during one. The Policy-Gated Tool Execution page is the technical half; the organizational half is who is allowed to say yes.
- Threat modelling belongs at design review, not before launch. A security review that arrives after the data boundary is set can only document risk, not remove it. Moving it earlier is a process change only a senior engineer can push through.
- Blast radius is the number to ask for. Not “is this secure” but “what can this component reach if it is fully compromised”. It is answerable, it is reviewable, and it converts a vague discussion into a scoping decision.
- Make the secure path the easy path. A reviewer catching secrets in code is a norm; a scanner failing the build is a mechanism. The same asymmetry as everywhere else in this module.
- Incidents are the highest-signal teaching moment and the worst time to assign blame. A review that produces a mechanism is worth more than one that produces a culprit, and the difference is visible to everyone watching how you run it.
Performance
Section titled “Performance”- Time-to-decision is a measurable quantity and usually a worse bottleneck than anyone admits. If a design decision sits for three weeks, that is three weeks of a team’s throughput, and it does not appear on any dashboard.
- Written and asynchronous beats a meeting for anything that needs consideration. It scales to people who were not there, it produces the artifact for free, and it surfaces objections from people who would not interrupt.
- Review latency compounds. A pull request waiting a day costs more than a day, because context has to be rebuilt on both sides. Fast review is a leadership behavior with a measurable effect on delivery.
- Count the meeting. Eight people for an hour is a working day. Some decisions are worth a working day; most recurring meetings have never been checked against that bar.
- Optimize for the team’s throughput, not your own. The senior engineer with the highest personal output is often the one whose review queue is the constraint on everyone else.
Scaling
Section titled “Scaling”- Mechanisms scale, attention does not. Every check that runs without you is capacity returned. This is the whole game past a certain team size.
- Delegate decision rights, not just tasks. “Own this component, including what goes in it” is transferable; “do this task” needs you again next week.
- Written artifacts are how you scale across time, not just headcount. The reader you are writing for is the engineer joining in a year with none of the context and no one left to ask.
- Pick the small set of things that must be consistent as the org grows and defend those hard — identity, observability, deployment. Defending everything means defending nothing.
- Grow the people who will replace you on the current work. The measure of technical leadership at this level is what keeps working after you move to the next problem, which is also the only way you ever get to move to the next problem.
Interview Questions
Section titled “Interview Questions”How do you drive a technical decision across teams that don't report to you?
Establish the constraints in writing first — most cross-team disagreement is about unstated requirements, not architecture. Then separate constraint disagreements from conclusion disagreements from preferences, and handle each differently: resolve constraints, buy evidence for conclusions, and let preferences go to whoever owns the code. Record the decision including what was rejected, then encode it in something that fails without you.
Tell me about a time you were wrong about a technical decision.
The answer that lands names what the evidence was, when it arrived, how long the gap was between the evidence and the reversal, and what changed so the next one is caught earlier. The gap is the interesting number — everyone has been wrong, and few have measured how long they took to notice.
How do you decide what to standardize and what to leave to teams?
Standardize where inconsistency has a compounding operational cost — identity, observability, deployment — because those are what on-call, incident response, and engineer mobility all depend on. Leave the rest, and be explicit that you are leaving it. The failure mode is standardizing on taste, which spends authority on something that produces no operational return.
How do you handle a disagreement with a senior colleague who has a different design?
Find out whether you disagree about constraints or conclusions. If constraints, resolve that first and the design argument often dissolves. If conclusions, name the evidence that would settle it and what it costs to get; a timeboxed spike is almost always cheaper than the argument. If it is preference within the same constraints, it is a two-way door — concede it and spend the credibility somewhere it matters.
What 'encode it in a mechanism' sounds like concretely
“We kept re-deciding whether a page needed its trade-offs written down, so I moved it into the content contract: a required section, checked in CI, failing the build with the file and the missing heading. It caught my own page the week after. The rule stopped needing a person, and the cost was about thirty lines.”
A model update improves benchmarks but breaks a downstream product assumption. How do you handle it?
This is a contract question wearing a model-quality costume. The immediate move is finding out whether the platform ever promised stability on that dimension — usually it did not, and nobody wrote it down. Short term, give the product team a pinned version and a migration window. Long term, the fix is a written contract for what the platform guarantees across model changes, plus an eval set the product team owns, so the next update surfaces the break before it ships rather than after.
Hands-on Lab
Section titled “Hands-on Lab”This module has no companion lab — its subject is the decisions around the code rather than the code. The closest thing to a lab is the ADR section, which is this repository practising what the module argues for, including the decisions that were later superseded.
Convert one norm into a mechanism
Find something your team says more than once a month — a review-checklist item, a naming rule, a “don’t forget to”. Write the check that makes forgetting it fail: a lint rule, a CI step, a required template section. Then verify it fails on a deliberately broken input, because an unfired check is indistinguishable from a broken one. Note how long it took; it is usually under an hour, which is the point.
Write the ADR for a decision already made
Take a significant decision your team made in the last year and write it up in the seven-section format above — Status, Context, Problem, Options, Decision, Consequences, References. The revealing section is Options: if you cannot reconstruct what was rejected and why, that decision is currently unauditable, and so is every one like it.
Signals you are operating at this level
- Decisions you are not in the room for still go the way the written record implies.
- The things you used to repeat are now checks that fail without you.
- You can name the last time you were wrong, and how long it took you to notice.
- Teams own the reversible decisions in their area outright, and know that they do.
- The irreversible decisions in your area have a written record naming what was rejected.
- Model or retrieval quality has a named owner and a threshold, not just latency.
- Someone else can do the work you were doing a year ago.
References
Section titled “References”- Google, Site Reliability Engineering, chapters on error budgets and postmortems — the reference worked example of replacing a recurring argument with a shared number.
- Will Larson, An Elegant Puzzle and Staff Engineer — the clearest writing on senior technical scope and the shape of the work at this level.
- Camille Fournier, The Manager’s Path — the chapters on tech lead and architect scope, useful precisely because they draw the line between technical leadership and management.
- Michael Nygard, Documenting Architecture Decisions — the original ADR format this handbook’s ADR section is built on.
- Module 0: Principal Engineer Mindset and Module 13: System Design — the framing and the method this module operates on.
Revision History
Section titled “Revision History”| Version | Date | Change |
|---|---|---|
| 1.0.0 | 2026-08-10 | Initial publication, completing the Learn track. |