Roadmap
Current version: v0.11.0 — The audio series
Section titled “Current version: v0.11.0 — The audio series”The documentation platform (v0.1.0) and all sixteen Learn modules are complete: Module 0 · Module 1 · Module 2 · Module 3 · Module 4: AI Infrastructure · Module 5: Agent Engineering · Module 6: MCP · Module 7: LangGraph · Module 8: RAG · Module 9: Model Serving · Module 10: Kubernetes · Module 11: Cloud · Module 12: Observability · Module 13: System Design · Module 14: Leadership · Module 15: Agent Identity. The Architecture section has a design-review page for all thirteen labs, each with a running implementation behind it, and the thirteenth — Multi-Region AI Serving Failover — shipped on the same day as its lab rather than after it. Build documents every lab from a run-it-and-verify-it angle. The Interview section now covers all three rounds: AI Infrastructure System Design, AI Systems Coding, and Technical Leadership. Reference is complete: thirteen lookups against a five-section page contract that CI enforces, with the four fast-moving ones — MCP, LangGraph, RAG, and Vector DB — each declaring the release it was verified against, on a 90-day review window CI fails once a page exceeds. Cheat Sheets is complete: five activity-scoped sheets against a four-section contract CI enforces — three for interview rounds, two for the work itself.
Every one of those pages can also be listened to. The audio series covers all 71 content pages — nearly twelve hours in total.
What’s next
Section titled “What’s next”| Version | Scope |
|---|---|
v0.2 |
|
v0.3 |
|
v0.4 |
|
v0.5 |
|
v0.6 |
|
v0.7 |
|
v0.8 |
|
v0.9 |
freshness block |
v0.10 |
|
v0.11 |
|
v0.12 |
|
v1.0 |
v1.0: shipped
Section titled “v1.0: shipped”Every criterion holds as of 2026-10-08:
- Architecture coverage is done again. v0.12 closed the gap, a twelfth lab reopened it, and Stateful Graph Checkpointing closes it on 2026-09-11 — twelve labs, twelve design reviews. The criterion has now been true, false, and true again, which is what it looks like when the closing condition is “every lab” and labs keep shipping.
- Interview coverage is done. All sixteen modules carry embedded questions, and all three rounds have a track.
- Every lab-shaped module has a lab. The criterion used to say every module has a lab, which could only be met by building three labs nobody should run. It now says every lab-shaped module has one. Modules 0, 13, and 14 are excluded by nature, not deferred: a lab about the Principal Engineer mindset, a system design round, or leadership would be a lab in name only, and the exclusion is stated here so it reads as a decision rather than an omission. The last two lab-shaped modules got theirs in turn: Module 7 on 2026-08-30 and Module 11 on 2026-10-08.
The architecture criterion held through the last lab for a reason worth keeping: the design review was written alongside the lab instead of after it, so the thirteenth lab never had a day without one. That is the practice to keep, because “every lab” is a closing condition that every new lab reopens.
After v1.0
Section titled “After v1.0”v1.0 is a floor, not a finish. Two things are already scheduled:
Re-verification of eleven fast-moving pages— done 2026-10-08, a month before their window closed. Each was checked against the current artifact it names rather than re-stamped, and it was not a formality: it found the MCP lab’s tests running the legacy protocol, its tokens accepted by any server, an SDK claim wrong since publication, a CIMD revision two behind, the LangGraph0.4line’s end of support in December 2026, and Pinecone’s API a quarter on. The LangGraph deadlock is still present in1.2.14. Next window closes 2027-01-06.- Decision models — the category Jev, Laya, OpenAI’s Decisions API, and Cloudflare’s Clef opened in September 2026: models that answer typed questions with probabilities instead of generating text. Planned as a Reference lookup, sections in Modules 4 and 5, a lab measuring calibration and the escalation threshold, and a design review. Nothing of it exists yet.
Labs are tracked separately from module versions, since a production-quality lab is a much larger unit of work than a module and shouldn’t gate a version bump the way a module does. Nothing on this table is planned — every lab below is built:
| Lab | Companion module | Status |
|---|---|---|
| Async AI Gateway | Module 1, 3 | Production-ready |
| SLO-Driven AI Operations | Module 10, 12 | Production-shaped |
| Durable Agent Task Engine | Module 2, 5 | Production-shaped |
| Policy-Gated Tool Runtime | Module 5 | Production-shaped |
| Hybrid Retrieval and Evaluation | Module 8 | Production-shaped |
| Dynamic Batching Inference | Module 9 | Production-shaped |
| Multi-Tenant MCP Server | Module 6 | Production-shaped |
| Agent Identity Broker | Module 15, 6 | Production-shaped |
| Model Router | Module 4 | Production-shaped |
| Semantic Cache | Module 4 | Production-shaped |
| Evaluation Platform | Module 4, 12 | Production-shaped |
| LangGraph Checkpoint Cost | Module 7 | Production-shaped |
| Region Failover Budget | Module 11 | Production-shaped |
Region Failover Budget: shipped with its design review
Section titled “Region Failover Budget: shipped with its design review”Module 11 was the last lab-shaped module with no lab. Its lab,
Region Failover Budget, measures the module’s own
FailoverController — and found it wrong in two ways the module did not say. A window tied to the
RTO spends the whole RTO on detection, so recovery overruns on every standby tier, hot included;
and a controller that returns False when it has no checks reads silence as health, so a region
that goes dark is never failed over — caught in 2 of 30 simulated runs, both by luck. The module
now carries a correction next to the code rather than a rewrite of it, so the measurement stays
reproducible.
It is a seeded simulation, not a cloud deployment, and says so; the downstream stages of a failover are inputs for the reader to replace. RPO is deliberately out of scope, for reasons the lab states.
LangGraph Checkpoint Cost: shipped, and now complete
Section titled “LangGraph Checkpoint Cost: shipped, and now complete”Module 7 was one of two lab-shaped modules with no lab. It has one
now: LangGraph Checkpoint Cost, built against a pinned
langgraph==1.2.11 rather than against the API from memory.
It measures what the module and the LangGraph lookup both assert
without a number. A plain accumulating channel writes about 1.34 MB over 100 steps; the same
accumulation through DeltaChannel writes about 60 KB — quadratic against linear. It also
measures the read side that buys, records that DeltaChannel is beta with an explicitly unstable
on-disk contract, and characterises a real deadlock in the pinned release’s DeltaChannel write
path that hung an unpatched 100-step sweep until it was killed.
It shipped without two things, and both have since been supplied:
- It has an Architecture page now. Stateful Graph Checkpointing was published on 2026-09-11 — the design review this lab shipped without, covering the snapshot-versus-delta trade, the replay depth a periodic snapshot bounds, and the batching-invariance condition the runtime cannot check for you.
- It has an episode now, and so does its design review. They were the 68th and 69th content pages, and the first two without audio since the series shipped. Both were recorded on 2026-10-08.
The audio series: shipped
Section titled “The audio series: shipped”Seventy-one episodes, nearly twelve hours of audio, from 3:00 on the shortest ADR to 22:24 on Module 6 — measured from the episode manifests, not estimated. The model plans the arc and writes the two-voice dialogue; a local text-to-speech model speaks it.
The measured total is worth stating plainly because
docs/PODCAST.md
projected seven, from the handbook’s 414 minutes of prose. Spoken dialogue expands on the page it
covers rather than reading it aloud, and the estimate did not account for that. The number here is
the measured one.
That is every page that says something. The ten pages that carry no episode by design — Start Here, this roadmap, and the eight section indexes — are navigation, and an episode narrating a table of contents would be filler. That is a deliberate exclusion, not a gap.
The four v0.12 architecture pages were recorded on 2026-08-28, four days after the rest. The recordings are worth a note of their own, because the run measured the estimator: two episodes landed within 1% of their quoted expectation and two came in 8% and 28% under it. The engine’s stated accuracy is about 7%, which is holding on the expensive side and not the cheap one. Every episode also overran its target length — 15:30 to 16:56 against targets near 12:00 — exactly as the engine’s own documentation warns that models overrun the character budget.
The last two, recorded on 2026-10-08, repeated both patterns. They came in 10% and 12% under their quoted expectation ($0.51 against $0.57, $0.33 against $0.37), so the estimator still over-quotes rather than under-quotes. And both overran: 17:05 against a 15:05 target, and 14:28 against 9:54 — the second by nearly half, on a lab page dense with identifiers that the dialogue explains rather than reads.
Two things about the pipeline are recorded rather than assumed. Speech runs locally, so the only
money the series spends is on the language model; and every transcript is committed next to its
manifest, so re-rendering a page’s episode text never pays the model twice.
ADR-0008 covers why the engine is built against
two ports — LlmPort and TtsPort — with no vendor SDK imported anywhere behind them, which is
what lets the whole pipeline be tested without a network or an API key.
Audio is served from a separate origin and is not in the repository; the transcripts are, so the words survive independently of the hosting.
Module 15 — Agent Identity and Access: shipped
Section titled “Module 15 — Agent Identity and Access: shipped”Suggested by Shantanu Lodh, and the gap is real: the handbook covers what an agent is allowed to do without covering who it is. Module 5 and the Policy-Gated Tool Runtime enforce capability scopes and approvals, and Module 6 argues that authorization belongs on the transport rather than in per-request metadata — but none of them answer what identity the transport is carrying.
The module is scoped around two questions:
- Agent-scoped credentials versus blanket delegation. Treating every agent action as on-behalf-of the user is the default because it is the easiest thing to build, and it makes the agent’s blast radius equal to the user’s entire permission set. Short-lived credentials minted per agent — narrowed in audience, scope, and lifetime — bound the damage, make actions attributable to the agent rather than the human, and make revocation something other than all-or-nothing.
- Authorization for remote MCP servers. A remote server has to decide whether a caller is allowed to invoke a tool, and the answer cannot be “the client said so”. That means the host registering with an identity provider, users authenticating through OIDC/OAuth, and the server validating audience and scope server-side on every call — the same argument Module 6 makes about the transport, taken one layer down into what the token actually asserts.
Module 15 is published, written against the MCP 2026-07-28
authorization specification and the OAuth RFCs beneath it — token exchange (RFC 8693), resource
indicators (RFC 8707), and the audience validation the specification requires of every MCP server.
The companion lab, Agent Identity Broker, is built. It mints scoped short-lived tokens, validates audience and scope server-side, and measures the blast-radius difference across a three-server fleet — one tool opened by a fully narrowed token, two when only the audience is narrowed. CI runs a mutation job that disables audience verification and fails the build if the suite stays green.
Lab migration: complete
Section titled “Lab migration: complete”A parallel branch built six labs against the retired static-HTML prototype before this platform
existed. All six now live in labs/, each re-verified and honestly re-labelled on arrival rather
than moved wholesale. That branch has been deleted; nothing on it is unique. The other five labs —
the Async AI Gateway and the four built in August 2026 for agent
identity, model routing, semantic caching, and evaluation — were written on this platform and never
migrated.
Five were straightforward ports: SLO-Driven AI Operations, Durable Agent Task Engine, Policy-Gated Tool Runtime, Hybrid Retrieval and Evaluation, and Dynamic Batching Inference.
The Multi-Tenant MCP Server was rebuilt rather than
migrated. The donor version targeted 2025-11-25 and described itself as “MCP-style”, so porting
it would have imported a protocol that no longer exists. It is now a real MCP server on the official
SDK (mcp 2.0.0, which reports 2026-07-28 as its only modern protocol version), tested through
the SDK’s own client over Streamable HTTP.
That rebuild took two attempts, and the failed one is recorded because it is the more useful half:
a credential carried in per-request _meta cannot secure an MCP server. The SDK issues protocol
calls of its own — call_tool() internally calls validate_tool_result(), which issues its own
tools/list. Both requests carry the SDK’s protocol _meta stamp, but only the first can carry the
application’s: call_tool() takes a meta= argument and list_tools() has none. So a server
authorizing on _meta rejects its own client, and a server that exempts tools/list to compensate
has reopened the hole. Authorization sits on the
transport, where every request carries it. See
Module 6 and the lab’s own write-up.
The Async AI Gateway needed no migration but did need a fix: it
was labelled production-ready while its CI omitted a dependency extra, so test collection aborted
before anything ran — masking 13 type errors and a failing test. A label is a claim, and that one
was not being checked.
How this roadmap works
Section titled “How this roadmap works”Each version bump ships complete modules, not partial ones — a module isn’t “in progress” on this
roadmap, it’s either published (with every required section from the Learn structure)
or not started. scripts/lint-content-structure.ts enforces that a merged module page has every
required section, so the roadmap can’t silently drift from what’s actually on the site. Labs follow
their own schedule in the table above rather than gating a module’s version bump — a module ships
once its content is complete, whether or not its companion lab exists yet.
Fast-moving pages carry a standing obligation the roadmap does not otherwise track. Each declares
verifiedAgainst and verifiedOn, and CI fails the build once the page passes its 90-day review
window — so the next re-verification pass is scheduled by the linter rather than by anyone
remembering. The pages published in v0.9 come due in November 2026. Two of them —
RAG and Vector DB — cover subject matter
with no version of its own, so each names a real implementation baseline (LangChain and Pinecone’s
dated API) instead. A fabricated version number would satisfy the linter and mean nothing.
Track progress against this roadmap through the repository’s releases and commit history.