Start here
The handbook covers a small number of systems in depth, and looks at each one from four angles: the concepts behind it, a design review of it, running code for it, and the interview round where it comes up. Those angles are the sections in the sidebar. This page explains which one to open, and when.
Who this is for, and what it is not
Section titled “Who this is for, and what it is not”It is written for engineers who already build with models and are working toward — or interviewing for — Principal and Staff AI Engineer roles. It assumes ML fundamentals rather than teaching them: no page here derives backpropagation or explains how attention works. What it covers instead is the layer above: how to design, run, and defend a production AI system, and how that gets examined in a Principal-level loop.
That scope shapes three choices worth knowing before you start:
- Depth on a few systems, not breadth across the field. Thirteen systems, each seen from four angles, rather than a linear course from first principles. If you want to know whether the handbook covers a topic, the table below is the honest answer.
- The code is production-shaped, not from-scratch. The labs are services — typed, tested, and gated in CI — that measure a trade-off, not reimplementations of an algorithm to show how it works. “From scratch” here would mean something different and less useful.
- The handbook is itself a worked example. Its Decision Records document real trade-offs made building it, and they are written in the same form the Architecture pages ask you to write yours.
If it is the fundamentals you need, learn them first and come back. AI Engineering from Scratch is a free, open-source curriculum that builds the core algorithms by hand, from linear algebra up, in 523 lessons across twenty phases. The two cover different ground and fit in sequence: that one for how models work, this one for the systems built around them.
The sections
Section titled “The sections”| Section | What it is | Open it when |
|---|---|---|
| Learn | Sixteen modules, each with the same fifteen sections — mental model, architecture, implementation, failure modes, trade-offs, interview questions | You want to understand a topic properly, start to finish |
| Build | Thirteen labs, each a real Python service with tests, type checking, and CI | You want to run something, break it, and see what happens |
| Architecture | Thirteen design reviews — constraints, request flow, failure modes, cost, observability, deployment | You are designing something similar and need the decisions laid out |
| Interview | Tracks organized by interview round, routing into the material that answers each one | You have an interview and want to know what the round samples |
| Reference | One-page lookups: what it is, the numbers, the gotchas | You need to check one fact quickly |
| Cheat Sheets | Printable one-page summaries | You want something to review on paper |
| Decision Records | Why the handbook and its labs are built the way they are | You are curious how a decision here was reached, or want the format |
Every section opens with an Overview that indexes it. The sidebar collapses to one line per section, and expands whichever section you are currently reading — so the open group is always the answer to “where am I”.
How one topic runs across the sections
Section titled “How one topic runs across the sections”This is the structure worth knowing. A topic is not confined to one section: it appears as a module, a design review, a lab, and a set of interview questions, and those four are cross-linked.
Read a row left to right and you have gone from the idea to a design you could defend to code you can run. Read it right to left when you already have the code and need the reasoning behind it.
Three routes through the material
Section titled “Three routes through the material”You have an interview coming up. Start at Interview and pick the track for the round you are facing. Each track says what that round actually samples and links to the material that answers it. Interview questions themselves live inside the modules and architecture pages, next to the material that makes the answers credible — the tracks route you there rather than copying them out.
You want to build something. Start at Build, pick the lab closest to your problem, and run it. Each lab page says what is real and what is simulated before you invest time. Then read the matching architecture page for the decisions behind it, using the table above.
You are designing a system at work. Start at the Architecture page for the closest system and read its Constraints and Failure Modes sections first — those are the two that transfer. Module 13 covers the method itself: how to establish constraints, model capacity, and identify which decision is hardest to reverse.
If none of those fit, Module 0 is the intended front door and reads as a standalone piece.
What the labels mean
Section titled “What the labels mean”The handbook makes claims about its own material, and they are meant literally.
- Lab maturity.
production-shapedmeans the architecture, tests, and failure handling are real but something a deployment needs is deliberately simulated — an in-memory store, a deterministic fake provider.production-readymeans the remaining stand-ins have a real integration path. Each lab page names its stand-ins explicitly. - Freshness. Pages on fast-moving topics carry what they were verified against and when. CI fails once such a page passes its review window, so a stale page is a build failure rather than something you discover by trusting it.
- Page contracts. Every module has the same fifteen sections, every architecture page the same twelve, every lookup the same five. CI enforces it. Once you know the shape of one page you can navigate all of them, and jump straight to the section you need.
Conventions worth knowing
Section titled “Conventions worth knowing”- Interview questions are never separated from their material. They sit at the bottom of the module or architecture page that explains the answer.
- Links go to sections, not just pages. A link to
#trade-offsmeans the trade-offs section specifically, and it is worth following. - Code snippets are taken from the labs, not written for the page. If a snippet looks incomplete, the lab has the rest.
- The Roadmap is honest about what is not written yet. Anything listed as planned does not exist.