Skip to content

Roadmap

Current version: v0.11.0 — The audio series

Section titled “Current version: v0.11.0 — The audio series”

The documentation platform (v0.1.0) and all sixteen Learn modules are complete: Module 0 · Module 1 · Module 2 · Module 3 · Module 4: AI Infrastructure · Module 5: Agent Engineering · Module 6: MCP · Module 7: LangGraph · Module 8: RAG · Module 9: Model Serving · Module 10: Kubernetes · Module 11: Cloud · Module 12: Observability · Module 13: System Design · Module 14: Leadership · Module 15: Agent Identity. The Architecture section has a design-review page for all thirteen labs, each with a running implementation behind it, and the thirteenth — Multi-Region AI Serving Failover — shipped on the same day as its lab rather than after it. Build documents every lab from a run-it-and-verify-it angle. The Interview section now covers all three rounds: AI Infrastructure System Design, AI Systems Coding, and Technical Leadership. Reference is complete: thirteen lookups against a five-section page contract that CI enforces, with the four fast-moving ones — MCP, LangGraph, RAG, and Vector DB — each declaring the release it was verified against, on a 90-day review window CI fails once a page exceeds. Cheat Sheets is complete: five activity-scoped sheets against a four-section contract CI enforces — three for interview rounds, two for the work itself.

Every one of those pages can also be listened to. The audio series covers all 71 content pages — nearly twelve hours in total.

Version Scope
v0.2 Learn Module 0 (Principal Engineer Mindset) and Module 1 (Production Python) — shipped 2026-08-05
v0.3 Learn Modules 2–3 (Distributed Systems, Networking) and the first Architecture page — shipped 2026-08-05
v0.4 Learn Modules 4–7 (AI Infrastructure, Agent Engineering, MCP, LangGraph) — shipped 2026-08-05
v0.5 Learn Modules 8–9 (RAG, Model Serving) — shipped 2026-08-07
v0.6 Learn Modules 10–11 (Kubernetes, Cloud) — shipped 2026-08-07
v0.7 Learn Module 12 (Observability) — shipped 2026-08-07
v0.8 Learn Module 14 (Leadership) — shipped 2026-08-10
v0.9 The four fast-moving Reference lookups (MCP, LangGraph, RAG, Vector DB), each with a freshness block — shipped 2026-08-12
v0.10 Learn Module 15 (Agent Identity and Access) — agent-scoped credentials and remote MCP authorization — shipped 2026-08-14
v0.11 The audio series — a generated episode on every content page, with its transcript — shipped 2026-08-23
v0.12 The four missing design reviews — agent identity, model routing, semantic caching, model evaluation — shipped 2026-08-28
v1.0 Complete, cross-linked handbook: every Learn module has an Architecture page and interview coverage, and every lab-shaped module has a lab — shipped 2026-10-08

Every criterion holds as of 2026-10-08:

  • Architecture coverage is done again. v0.12 closed the gap, a twelfth lab reopened it, and Stateful Graph Checkpointing closes it on 2026-09-11 — twelve labs, twelve design reviews. The criterion has now been true, false, and true again, which is what it looks like when the closing condition is “every lab” and labs keep shipping.
  • Interview coverage is done. All sixteen modules carry embedded questions, and all three rounds have a track.
  • Every lab-shaped module has a lab. The criterion used to say every module has a lab, which could only be met by building three labs nobody should run. It now says every lab-shaped module has one. Modules 0, 13, and 14 are excluded by nature, not deferred: a lab about the Principal Engineer mindset, a system design round, or leadership would be a lab in name only, and the exclusion is stated here so it reads as a decision rather than an omission. The last two lab-shaped modules got theirs in turn: Module 7 on 2026-08-30 and Module 11 on 2026-10-08.

The architecture criterion held through the last lab for a reason worth keeping: the design review was written alongside the lab instead of after it, so the thirteenth lab never had a day without one. That is the practice to keep, because “every lab” is a closing condition that every new lab reopens.

v1.0 is a floor, not a finish. Two things are already scheduled:

  • Re-verification of eleven fast-moving pages — done 2026-10-08, a month before their window closed. Each was checked against the current artifact it names rather than re-stamped, and it was not a formality: it found the MCP lab’s tests running the legacy protocol, its tokens accepted by any server, an SDK claim wrong since publication, a CIMD revision two behind, the LangGraph 0.4 line’s end of support in December 2026, and Pinecone’s API a quarter on. The LangGraph deadlock is still present in 1.2.14. Next window closes 2027-01-06.
  • Decision models — the category Jev, Laya, OpenAI’s Decisions API, and Cloudflare’s Clef opened in September 2026: models that answer typed questions with probabilities instead of generating text. Planned as a Reference lookup, sections in Modules 4 and 5, a lab measuring calibration and the escalation threshold, and a design review. Nothing of it exists yet.

Labs are tracked separately from module versions, since a production-quality lab is a much larger unit of work than a module and shouldn’t gate a version bump the way a module does. Nothing on this table is planned — every lab below is built:

Lab Companion module Status
Async AI Gateway Module 1, 3 Production-ready
SLO-Driven AI Operations Module 10, 12 Production-shaped
Durable Agent Task Engine Module 2, 5 Production-shaped
Policy-Gated Tool Runtime Module 5 Production-shaped
Hybrid Retrieval and Evaluation Module 8 Production-shaped
Dynamic Batching Inference Module 9 Production-shaped
Multi-Tenant MCP Server Module 6 Production-shaped
Agent Identity Broker Module 15, 6 Production-shaped
Model Router Module 4 Production-shaped
Semantic Cache Module 4 Production-shaped
Evaluation Platform Module 4, 12 Production-shaped
LangGraph Checkpoint Cost Module 7 Production-shaped
Region Failover Budget Module 11 Production-shaped

Region Failover Budget: shipped with its design review

Section titled “Region Failover Budget: shipped with its design review”

Module 11 was the last lab-shaped module with no lab. Its lab, Region Failover Budget, measures the module’s own FailoverController — and found it wrong in two ways the module did not say. A window tied to the RTO spends the whole RTO on detection, so recovery overruns on every standby tier, hot included; and a controller that returns False when it has no checks reads silence as health, so a region that goes dark is never failed over — caught in 2 of 30 simulated runs, both by luck. The module now carries a correction next to the code rather than a rewrite of it, so the measurement stays reproducible.

It is a seeded simulation, not a cloud deployment, and says so; the downstream stages of a failover are inputs for the reader to replace. RPO is deliberately out of scope, for reasons the lab states.

LangGraph Checkpoint Cost: shipped, and now complete

Section titled “LangGraph Checkpoint Cost: shipped, and now complete”

Module 7 was one of two lab-shaped modules with no lab. It has one now: LangGraph Checkpoint Cost, built against a pinned langgraph==1.2.11 rather than against the API from memory.

It measures what the module and the LangGraph lookup both assert without a number. A plain accumulating channel writes about 1.34 MB over 100 steps; the same accumulation through DeltaChannel writes about 60 KB — quadratic against linear. It also measures the read side that buys, records that DeltaChannel is beta with an explicitly unstable on-disk contract, and characterises a real deadlock in the pinned release’s DeltaChannel write path that hung an unpatched 100-step sweep until it was killed.

It shipped without two things, and both have since been supplied:

  • It has an Architecture page now. Stateful Graph Checkpointing was published on 2026-09-11 — the design review this lab shipped without, covering the snapshot-versus-delta trade, the replay depth a periodic snapshot bounds, and the batching-invariance condition the runtime cannot check for you.
  • It has an episode now, and so does its design review. They were the 68th and 69th content pages, and the first two without audio since the series shipped. Both were recorded on 2026-10-08.

Seventy-one episodes, nearly twelve hours of audio, from 3:00 on the shortest ADR to 22:24 on Module 6 — measured from the episode manifests, not estimated. The model plans the arc and writes the two-voice dialogue; a local text-to-speech model speaks it.

The measured total is worth stating plainly because docs/PODCAST.md projected seven, from the handbook’s 414 minutes of prose. Spoken dialogue expands on the page it covers rather than reading it aloud, and the estimate did not account for that. The number here is the measured one.

That is every page that says something. The ten pages that carry no episode by design — Start Here, this roadmap, and the eight section indexes — are navigation, and an episode narrating a table of contents would be filler. That is a deliberate exclusion, not a gap.

The four v0.12 architecture pages were recorded on 2026-08-28, four days after the rest. The recordings are worth a note of their own, because the run measured the estimator: two episodes landed within 1% of their quoted expectation and two came in 8% and 28% under it. The engine’s stated accuracy is about 7%, which is holding on the expensive side and not the cheap one. Every episode also overran its target length — 15:30 to 16:56 against targets near 12:00 — exactly as the engine’s own documentation warns that models overrun the character budget.

The last two, recorded on 2026-10-08, repeated both patterns. They came in 10% and 12% under their quoted expectation ($0.51 against $0.57, $0.33 against $0.37), so the estimator still over-quotes rather than under-quotes. And both overran: 17:05 against a 15:05 target, and 14:28 against 9:54 — the second by nearly half, on a lab page dense with identifiers that the dialogue explains rather than reads.

Two things about the pipeline are recorded rather than assumed. Speech runs locally, so the only money the series spends is on the language model; and every transcript is committed next to its manifest, so re-rendering a page’s episode text never pays the model twice. ADR-0008 covers why the engine is built against two ports — LlmPort and TtsPort — with no vendor SDK imported anywhere behind them, which is what lets the whole pipeline be tested without a network or an API key.

Audio is served from a separate origin and is not in the repository; the transcripts are, so the words survive independently of the hosting.

Module 15 — Agent Identity and Access: shipped

Section titled “Module 15 — Agent Identity and Access: shipped”

Suggested by Shantanu Lodh, and the gap is real: the handbook covers what an agent is allowed to do without covering who it is. Module 5 and the Policy-Gated Tool Runtime enforce capability scopes and approvals, and Module 6 argues that authorization belongs on the transport rather than in per-request metadata — but none of them answer what identity the transport is carrying.

The module is scoped around two questions:

  • Agent-scoped credentials versus blanket delegation. Treating every agent action as on-behalf-of the user is the default because it is the easiest thing to build, and it makes the agent’s blast radius equal to the user’s entire permission set. Short-lived credentials minted per agent — narrowed in audience, scope, and lifetime — bound the damage, make actions attributable to the agent rather than the human, and make revocation something other than all-or-nothing.
  • Authorization for remote MCP servers. A remote server has to decide whether a caller is allowed to invoke a tool, and the answer cannot be “the client said so”. That means the host registering with an identity provider, users authenticating through OIDC/OAuth, and the server validating audience and scope server-side on every call — the same argument Module 6 makes about the transport, taken one layer down into what the token actually asserts.

Module 15 is published, written against the MCP 2026-07-28 authorization specification and the OAuth RFCs beneath it — token exchange (RFC 8693), resource indicators (RFC 8707), and the audience validation the specification requires of every MCP server.

The companion lab, Agent Identity Broker, is built. It mints scoped short-lived tokens, validates audience and scope server-side, and measures the blast-radius difference across a three-server fleet — one tool opened by a fully narrowed token, two when only the audience is narrowed. CI runs a mutation job that disables audience verification and fails the build if the suite stays green.

A parallel branch built six labs against the retired static-HTML prototype before this platform existed. All six now live in labs/, each re-verified and honestly re-labelled on arrival rather than moved wholesale. That branch has been deleted; nothing on it is unique. The other five labs — the Async AI Gateway and the four built in August 2026 for agent identity, model routing, semantic caching, and evaluation — were written on this platform and never migrated.

Five were straightforward ports: SLO-Driven AI Operations, Durable Agent Task Engine, Policy-Gated Tool Runtime, Hybrid Retrieval and Evaluation, and Dynamic Batching Inference.

The Multi-Tenant MCP Server was rebuilt rather than migrated. The donor version targeted 2025-11-25 and described itself as “MCP-style”, so porting it would have imported a protocol that no longer exists. It is now a real MCP server on the official SDK (mcp 2.0.0, which reports 2026-07-28 as its only modern protocol version), tested through the SDK’s own client over Streamable HTTP.

That rebuild took two attempts, and the failed one is recorded because it is the more useful half: a credential carried in per-request _meta cannot secure an MCP server. The SDK issues protocol calls of its own — call_tool() internally calls validate_tool_result(), which issues its own tools/list. Both requests carry the SDK’s protocol _meta stamp, but only the first can carry the application’s: call_tool() takes a meta= argument and list_tools() has none. So a server authorizing on _meta rejects its own client, and a server that exempts tools/list to compensate has reopened the hole. Authorization sits on the transport, where every request carries it. See Module 6 and the lab’s own write-up.

The Async AI Gateway needed no migration but did need a fix: it was labelled production-ready while its CI omitted a dependency extra, so test collection aborted before anything ran — masking 13 type errors and a failing test. A label is a claim, and that one was not being checked.

Each version bump ships complete modules, not partial ones — a module isn’t “in progress” on this roadmap, it’s either published (with every required section from the Learn structure) or not started. scripts/lint-content-structure.ts enforces that a merged module page has every required section, so the roadmap can’t silently drift from what’s actually on the site. Labs follow their own schedule in the table above rather than gating a module’s version bump — a module ships once its content is complete, whether or not its companion lab exists yet.

Fast-moving pages carry a standing obligation the roadmap does not otherwise track. Each declares verifiedAgainst and verifiedOn, and CI fails the build once the page passes its 90-day review window — so the next re-verification pass is scheduled by the linter rather than by anyone remembering. The pages published in v0.9 come due in November 2026. Two of them — RAG and Vector DB — cover subject matter with no version of its own, so each names a real implementation baseline (LangChain and Pinecone’s dated API) instead. A fabricated version number would satisfy the linter and mean nothing.

Track progress against this roadmap through the repository’s releases and commit history.