Track: AI Systems Coding
Read the transcript
1. What the round is actually grading
Host: So let’s start with what people get wrong before they even open the editor. They think this round is about naming the right pattern — ‘oh, I’d use a semaphore here, a worker pool there’ — and they treat that as basically the answer.
Guest: Right, naming it isn’t the answer. The real test is whether you can write a semaphore that actually gets released on every exit path, including when the request gets cancelled halfway through. It’s also whether you handle timeout, partial failure, and shutdown, not just the case where everything goes right. And it’s whether you picked the right concurrency model at all — async for something CPU-bound, or threads for work the GIL is just going to serialize anyway, that reads as not knowing the runtime, no matter how clean the code looks.
Host: And there’s this distinction you draw between a mid-level answer and a Principal answer — it’s not really about who writes more code, is it?
Guest: No, it’s almost the opposite. A mid-level candidate writes until the function works and stops. A Principal candidate states the invariant out loud first — something like ‘no more than N in flight, and a rejected request can’t consume a slot’ — writes the smallest thing that actually holds that invariant, and then names what they deliberately left out, like a bounded queue instead of an unbounded one, and what would force them to change that choice. Interviewers are grading that invariant and that omission as heavily as the code itself.
2. The three formats you’ll actually be handed
Host: So let’s map the terrain. Candidates walking into these loops actually run into three distinct formats, and they don’t all test the same thing. What are they?
Guest: There’s the applied build, where you’re handed something like ‘write a rate limiter’ and given forty-five to sixty minutes, usually in one file — that’s testing concurrency correctness and whether you handle cancellation and shutdown without being told to. There’s debug-and-extend, where you get an existing repo with a failing test, and that’s testing how you navigate unfamiliar code and form a hypothesis without breaking things you didn’t mean to touch. And there’s the algorithmic screen, the conventional data-structures problem, usually earlier in the loop — that’s just baseline fluency, still real, but not what later rounds actually weigh.
Host: And if someone only has time to prepare one of those well, which one should it be?
Guest: The applied build, no contest — it dominates Principal, Staff, AI Platform, and Founding AI Engineer loops because it’s the closest thing to a proxy for the actual job. And the best prep isn’t grinding more problems, it’s writing four or five of these components from an empty file, with tests, until the cancellation and shutdown paths are reflex instead of something you have to recall under pressure.
3. Ten components worth writing from an empty file
Host: Okay, four or five components — give me the actual list, because I think people’s mental model of what counts as a serious infra component is smaller than what interviewers actually reach for.
Guest: There’s a set of about ten that show up verbatim across loops: token-bucket rate limiter, retry with backoff and jitter, circuit breaker, bounded-concurrency worker pool, graceful drain, idempotency key store, lease-based checkout, a latency-bounded batcher, reciprocal rank fusion, and multi-window burn-rate alerting. Each one has a single invariant that the interviewer is actually listening for — not the code shape, the property that has to hold no matter what breaks.
Host: Give me one where the invariant is easy to say but the code trap is subtle — because I feel like the gateway itself has one baked in.
Guest: Right there in the admit path — rate limiting happens before the semaphore acquire, so a rejected caller never consumes concurrency capacity it was denied. And the whole retry loop sits inside one asyncio timeout block wrapping every attempt, so backoff can’t quietly blow past the caller’s deadline, with the semaphore released in a finally block no matter which branch exits. The drain pattern is the cleanest example of ‘finished’ — flip accepting to false, wait on a condition for active count to hit zero, and if the timeout fires, return false explicitly instead of silently killing requests mid-flight.
4. The traps that catch people in a predictable order
Host: Let’s go through the traps in the order they actually bite people. What’s number one, the thing that shows up before anyone even gets to concurrency logic?
Guest: Blocking the event loop. A time.sleep, a synchronous HTTP client, requests, a CPU-heavy loop dropped inside an async def — it freezes every coroutine sharing that loop, not just the slow one. Interviewers plant this on purpose, sometimes hand you code that already has it, and the fix is loop.run_in_executor or a separate process, never inline.
Host: And once someone’s past that, where’s the next one waiting?
Guest: Semaphore leaks — acquire before a try, forget the finally, or only release on the success branch, and capacity leaks until the service wedges, faster under cancellation because nobody tested that path. Right behind it is swallowing CancelledError, a bare except or even a correct except that doesn’t re-raise, which makes the task uncancellable. Then unbounded gather on caller-controlled input, sequential tests that prove nothing about the race you were asked to guarantee, and an in-process counter that’s correct on your laptop and wrong the moment it’s running on three replicas — fine as a starting point, only if you say so out loud.
5. Trade-offs you’ll be asked to defend mid-problem
Host: So beyond just not writing the bugs, the interviewer’s going to stop you mid-build and ask why you chose async over threads, or in-process over distributed. Where do those questions actually come from?
Guest: They come straight out of the same file. Take the rate limiter — it’s an in-process token bucket, which is fast and dependency-free, but it’s only correct on one replica. Run three replicas and a tenant effectively gets triple their quota, because each process enforces its own count with no idea the others exist. The distributed version, Redis-backed with an atomic Lua script, fixes that — but now you own a new failure mode: what does enforcement do when Redis is unavailable? The right answer in a 45-minute exercise is usually in-process, stated as such, with the distributed version described — same as async versus threads versus processes: async for I/O with async-native libraries, threads only when a blocking library forces your hand since the GIL kills any real CPU gain, processes when the work is actually CPU-bound and you can eat the serialization cost.
Host: And that same reject-versus-queue call shows up in the gateway itself, not just the limiter — so how do you keep from sounding like you’re dodging the question with ‘it depends’?
Guest: You make it depend on something specific, not on vibes. Rejecting at the door is a fast, honest error that protects the system; queueing or falling back keeps the request alive but can turn a bounded failure into an unbounded one, so you say which operation gets which and name the signal that decides. Same discipline on typing: full types and validation on the public signature and the data model, since that’s what the interviewer’s eyes are on and what stops the class of bug you’d otherwise chase live, but you don’t spend your remaining time annotating a private helper nobody’s going to see.
6. Debug-and-extend: navigation over authorship
Host: So now the flip side — instead of an empty file, they hand you a whole repo and a failing test. That feels like a totally different skill. Where do you even start?
Guest: You start by running the failing test before you read a line of code, because a symptom you’ve actually observed beats any theory you form from staring at source. Then you go find the invariant the code is trying to hold and the exact point where it stops holding — usually a missing release, an unawaited coroutine, an ordering assumption, or state leaking across a boundary that was supposed to isolate it. From there you change the smallest thing possible, because a fix that drags a refactor along with it is a fix nobody can review, and in an interview it reads as you not being able to scope your own change. Then you volunteer the regression test that would’ve caught it — that’s worth more than the fix itself — and you close by naming the adjacent thing you noticed and deliberately left alone, so it reads as judgment instead of something you missed.
Host: That last step feels like the one people skip, since it means admitting you saw a problem and walked past it.
Guest: Right, but the adjacent issue you noticed and deliberately left alone is a signal, not an omission, as long as you name it. And you can drill this whole sequence with the labs on your own — pick one, go to a scratch branch, deliberately break an invariant, and run the existing suite against it. Wherever the suite stays green, you’ve just found the exact test worth writing, which is the same muscle the round is testing.
7. Where to find the actual questions
Host: So if someone wants to drill the actual questions rather than just the labs, where do they go? Is it scattered across the handbook or is there a map?
Guest: It’s kept right next to the material that answers it, so you’re not hunting. Module 1 has event-loop behavior, cancellation, timeouts, backpressure — that’s the highest-yield set for this whole track. Module 2 covers idempotency and delivery guarantees under load, Module 9 has the batching latency-throughput trade, and Module 12 has the observability follow-up, because almost every build ends with someone asking what you’d actually measure.
8. The preparation checklist, and narrating while typing
Host: Okay, let’s land this with the actual checklist, because ‘study more’ isn’t a plan. What are the concrete things someone should be able to point to before they walk into this round?
Guest: Five things. One, you’ve written a token-bucket limiter, a retry-with-jitter wrapper, and a bounded worker pool from an empty file within the last month, and each has a test that fails when you break the invariant, not one that just checks the happy path. Two, you can explain what CancelledError does to a running task and where cleanup belongs, and you can say without hedging whether asyncio, threads, or processes fits a given workload. Three, you’ve read at least one lab’s tests end to end so you recognize what an invariant looks like as an assertion. Four, you can name what your solution doesn’t handle — replicas, persistence, ordering — before the interviewer has to ask. And five, you’ve practiced narrating while you type, because this round is scored on reasoning made audible, and ten minutes of silence loses points that a correct answer at the end doesn’t recover.
Host: That last one feels like the whole episode in one line — the code was never really the point, the thinking out loud was. That’s a great place to leave it, thanks for walking through all of this.
Not covered
The planner wanted these and found nothing in the source to support them:
- specific compensation or offer-negotiation guidance for these roles
- which named companies use each interview format
- how many total rounds make up a typical loop or how they’re sequenced
- guidance on take-home versus live-coding format preferences
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
The coding round for AI infrastructure roles is rarely a graph algorithm. It is usually a small, real concurrency or correctness problem — a limiter, a retry path, a batcher, a queue — that the interviewer has watched dozens of people fail in the same four or five ways. This track says what the round samples, lists the components worth being able to write from an empty file, and points at the lab source where each one exists finished, tested, and type-checked.
What is actually being assessed
Section titled “What is actually being assessed”The systems in Learn and Architecture are made of a small number of primitives, and the coding round checks whether you can actually build one rather than name it.
- Can you write correct concurrent code, not just describe it? The gap between “I would use a semaphore” and a semaphore released on every exit path — including cancellation — is the whole round.
- Do you handle the paths that aren’t the happy one? Timeout, cancellation, partial failure, and shutdown. Code that only works when nothing goes wrong is the single most common failing answer.
- Do you reach for the right concurrency model? Choosing async for CPU-bound work, or threads for something the GIL will serialize, reads as unfamiliarity with the runtime regardless of how clean the code is.
- Can you test what you wrote? Not “I would add tests” — an actual assertion that would fail if the property you claim were broken. Concurrency bugs are invisible to tests written by someone who hasn’t thought about which interleaving is the dangerous one.
- Do you know what your code costs? Memory growth under load, connection lifetime, and what happens at ten times the request rate.
At this level the code is a conversation, not a submission
A mid-level candidate writes until the function works. A Principal-level candidate states the invariant first (“no more than N in flight, and a rejected request must not consume a slot”), writes the smallest thing that holds it, then names what they left out and why — bounded queue instead of unbounded, in-process instead of distributed, and what would force the other choice. Interviewers grade the invariant and the omission at least as heavily as the code.
The three formats
Section titled “The three formats”| Format | What you are handed | What it is really testing |
|---|---|---|
| Applied build | “Write a rate limiter / retry wrapper / batching queue.” 45–60 min, usually one file | Concurrency correctness and whether you cover cancellation and shutdown without being prompted |
| Debug and extend | An existing repository with a failing test or a described symptom | How you navigate unfamiliar code, form a hypothesis, and avoid changing behavior you didn’t intend to change |
| Algorithmic screen | A conventional data-structures problem, usually earlier in the loop | Baseline fluency. Real, still filtered on, and not what the later rounds weigh |
The applied build dominates for Principal, Staff, AI Platform, and Founding AI Engineer loops, because it is the closest available proxy for the work. Prepare it first.
Engineering Note
The strongest preparation is not solving more problems — it is writing four or five of the components below from an empty file, with tests, until the cancellation and shutdown paths are reflex rather than recall. Every one of them appears in the labs, so you can compare your version against a finished one immediately after.
Components worth writing from an empty file
Section titled “Components worth writing from an empty file”Each of these has been asked, verbatim or nearly, in AI infrastructure loops. Each links to the module that explains the design and the lab file where it exists finished.
| Component | The invariant it must hold | Explained in | Finished in |
|---|---|---|---|
| Token-bucket rate limiter | Refill is time-based, not tick-based; a rejected caller consumes nothing | Module 1 | rate_limit.py |
| Retry with exponential backoff and jitter | Bounded attempts, bounded total deadline, jitter to break synchronization | Module 1 | gateway.py |
| Circuit breaker | State is derived from elapsed time, so no background timer is needed; one probe in half-open | Module 1 | resilience.py |
| Bounded-concurrency worker pool | The slot is released on success, failure, and cancellation alike | Module 1 | worker.py |
| Graceful drain on shutdown | Stop accepting first, then wait for in-flight work, then exit — with a timeout | Module 1 | draining.py |
| Idempotency key store | Claim-before-execute must be atomic, or two concurrent submissions both execute | Module 2 | store.py |
| Lease-based queue checkout | A lease must be fenced, so a stalled worker cannot complete work already reassigned | Module 2 | store.py |
| Latency-bounded request batcher | Flush on size or deadline, whichever comes first; a single request must not wait forever | Module 9 | batcher.py |
| Reciprocal rank fusion | Combines rankings, not scores — so two retrievers with incomparable scales still merge | Module 8 | fusion.py |
| Multi-window burn-rate alerting | Fast and slow windows must both fire before paging, or a blip pages someone | Module 12 | slo.py |
Read the corresponding tests as well as the source. In every lab, the test file is where the invariant is stated as an executable claim — which is exactly what the interviewer is listening for when they ask “how would you know this works?”
The traps, in the order they come up
Section titled “The traps, in the order they come up”Blocking the event loop inside async code
time.sleep(), a synchronous HTTP client, requests, or a CPU-heavy loop inside async def
freezes every coroutine on that loop — not just the one request. Interviewers plant this
deliberately; some hand you code that already contains it. Blocking work belongs in
loop.run_in_executor or a separate process. See
Module 1 → Failure Modes.
Acquiring a resource without releasing it on every path
A semaphore.acquire() before a try whose finally you forgot, or a release that only runs on
the success branch, leaks capacity until the service wedges. Under cancellation it leaks faster,
because the exception path is the one nobody tested. Either async with the primitive or put the
release in finally — and say out loud which you chose.
Swallowing CancelledError
A bare except Exception around an await catches cancellation in older code, and even a correct
except asyncio.CancelledError that does not re-raise makes the task uncancellable — a shutdown
that never completes and a deadline that never takes effect. Cancellation is control flow: clean
up in finally, then re-raise.
Unbounded fan-out
asyncio.gather() over a list whose length comes from the request is an easy way to exhaust
connections, provider quota, and memory in a single call. If the input size is caller-controlled,
the concurrency bound belongs inside the loop, not around it.
Testing concurrency by running it once
A test that calls the function sequentially proves nothing about the property you were asked to guarantee. The test that earns the point launches the concurrent calls that would race, then asserts the invariant — that the work executed exactly once, that no more than N ran at a time, that the rejected caller consumed no capacity.
Mutable state shared across replicas as if there were one
An in-process counter, cache, or token bucket is correct in the file you just wrote and wrong the moment it runs on three replicas — a tenant gets three times their limit. This is fine as a starting point; it is not fine unstated. Name it as a deliberate scope decision and describe what replacing it costs.
Trade-offs you will be asked to defend mid-problem
Section titled “Trade-offs you will be asked to defend mid-problem”asyncio, threads, or processes
Async wins for high-concurrency I/O with async-native libraries. Threads earn their keep only when a required library is blocking and has no async equivalent — they do not improve pure-Python CPU throughput, because of the GIL. Processes are the only real answer for CPU-bound Python, at the cost of serialization and per-process startup. See Module 1 → Trade-offs.
In-process state or a shared store
In-process is faster, has no extra dependency, and is correct on exactly one replica. A shared store makes the behavior correct under horizontal scaling and adds a new failure mode: what enforcement does when the store itself is unavailable. The right answer in a 45-minute exercise is usually in-process, stated as such, with the distributed version described.
Fail fast or degrade
Rejecting at the door gives callers a fast, honest error and protects the system. Queueing or falling back keeps requests alive and can turn a bounded failure into an unbounded one. The distinguishing answer picks per operation rather than globally, and names the signal that decides.
Typed and validated, or terse
Type annotations and validated boundaries catch a class of bug before runtime and make the code readable to the interviewer live. They also cost minutes you may not have. The workable line under time pressure: types on the public signature and the data model, nothing ceremonial inside a helper.
Debug-and-extend rounds
Section titled “Debug-and-extend rounds”When you are given a repository rather than an empty file, the round is testing navigation and restraint more than authorship.
- Reproduce before reading. Run the failing test first. A symptom you have observed is worth more than a hypothesis formed from reading.
- Find the invariant the code is trying to hold, then find where it stops holding. Most bugs at this level are a missing release, an unawaited coroutine, an ordering assumption, or state shared across a boundary that was supposed to isolate it.
- Change the smallest thing. A fix that also refactors is a fix nobody can review, and in an interview it reads as an inability to scope.
- Add the test that would have caught it. Volunteering this is worth more than the fix.
- Say what you did not change and why. The adjacent issue you noticed and deliberately left alone is a signal, not an omission — as long as you name it.
The labs are usable as practice material for exactly this: pick one, break an invariant on a scratch branch, and see whether the existing test suite catches it. Where it doesn’t, you have found a test worth writing — which is the same exercise the round is simulating.
Where the questions live
Section titled “Where the questions live”The handbook keeps interview questions next to the material that answers them. For this round:
- Module 1: Production Python — event-loop behavior, cancellation, timeouts, and backpressure. The highest-yield set for this track.
- Module 2: Distributed Systems — idempotency, delivery guarantees, and queue behavior under load.
- Module 9: Model Serving — batching and the latency/throughput trade it encodes.
- Module 12: Observability — what you measure, which is the follow-up question to almost every applied build.
Preparation checklist
Section titled “Preparation checklist”Before a coding round
- You have written a token-bucket limiter, a retry-with-jitter wrapper, and a bounded worker pool from an empty file, without looking, within the last month.
- Each of those has a test that fails if the invariant breaks — not a test that only proves the happy path returns.
- You can explain what
asyncio.CancelledErrordoes to a running task, and where cleanup belongs. - You can say, without hedging, which of asyncio, threads, and processes fits a given workload.
- You have read at least one lab’s tests end to end, so you know what stating an invariant as an assertion looks like.
- You can name what your solution does not handle — replicas, persistence, ordering — before the interviewer asks.
- You have practiced narrating while typing. The round is scored on reasoning you make audible, and silence for ten minutes loses points a correct answer does not recover.