Durable Agent Task Engine
Read the transcript
1. Why synchronous agent work breaks under interruption
Host: So let’s start with the failure everyone building agent systems eventually hits. You’ve got a multi-step run that dies on step seven, and if you’re not careful it restarts at step one, burning money and time it already spent. Or a client gets impatient, retries a submission, and now you’ve got two runs doing the same expensive work. Why does this keep happening even to teams who think they’ve handled it?
Guest: Because the default architecture is ordinary synchronous request handling, and that ties the work’s lifetime to the request’s lifetime. A deploy, a timeout, an OOM kill — any of those destroys progress that cost real money to produce. And the usual first fix, spinning up a background thread or firing off a detached task, doesn’t actually fix anything. It just hides the problem, because nothing durable records that the work exists, so when the worker dies there’s nothing left to recover it. What you actually need is a real delivery guarantee, and the honest one to build around is at-least-once — exactly-once delivery isn’t something you can have, you can only get exactly-once effect by making your handlers idempotent. Everything we’re going to talk about today, every mechanism, is really just downstream of accepting that one fact.
2. The visibility timeout as crash recovery, and why fencing is what makes it safe
Host: So walk me through the lease mechanism itself. When a worker picks up a task, what actually happens, and what makes it self-healing when that worker dies?
Guest: A worker leases the task and gets back two things: a fencing token and a visibility deadline. While that lease holds, the task is invisible to every other worker, so nobody double-picks it. If the worker dies — deploy, OOM, whatever — nobody renews the lease, it just expires on its own clock, and the task falls back to pending automatically. There’s no heartbeat service watching for that, no watchdog process, no liveness probe. The absence of a renewal is the only signal needed, which means there’s one fewer component in the system that can itself go down.
Host: Okay but that raises an obvious hole — what if the worker isn’t actually dead, just paused? A long GC pause, a suspended VM, something blocking on a syscall. It wakes back up still thinking it owns the task.
Guest: Right, that’s a zombie, and it’s exactly the case the lease alone doesn’t cover — the task may already have been reissued and finished by someone else while it was asleep. That’s what the fencing token is for. Every ack, fail, or checkpoint carries the token it was issued, and once the lease expires and a new lease is granted, that old token is stale and gets rejected outright. So the zombie wakes up, tries to mark its work done, and the system just refuses it instead of silently corrupting a task another worker already completed. That’s the difference between at-least-once being merely tolerable and actually being safe.
3. The subtle bug: counting failures vs counting deliveries
Host: Okay, so the fencing token stops the zombie from corrupting a finished task. But what stops the zombie’s *task* from just being retried forever? Isn’t that also a crash-driven loop the budget should catch?
Guest: That’s the bug that only shows up under fire, and it’s a subtle one. If your retry budget only decrements on an explicit failure — a handler catching an error and reporting it — then a handler that gets its whole process killed never reports anything. The lease just expires, the task goes back to pending with its budget completely untouched, and it gets redelivered forever. A task that reliably segfaults its worker becomes an infinite loop that also takes a worker down with it every single pass.
Host: So the fix is counting the attempt at delivery time, not waiting for someone to confess failure.
Guest: Exactly — decrement the budget the moment the task is handed out, and now crash-driven redelivery is bounded by the same limit as explicit failure. It dead-letters, an operator gets paged, the loop stops. And that’s a separate concern from idempotent submission, by the way — the idempotency key is there to stop a redelivered task from being executed twice, not to stop it from being created twice, but this is about a task that only got created once and still needs its execution attempts bounded no matter how it dies.
4. From lab to production: what the in-memory store stands in for
Host: So let’s zoom out for a second. Everything we’ve described so far — the lock, the lease table, the retry counter — that’s all living in one process’s memory in this lab. What breaks the moment you run more than one worker?
Guest: The single lock stops being correct, because it can only serialize claims within its own process — it has no idea a second replica exists. In production you need that same atomicity from the backing store itself, whether that’s row-level locking with SKIP LOCKED or an atomic Lua script: the point is contending workers step over a locked row instead of queueing behind it or, worse, both grabbing it. Same discipline extends outward too — graceful draining so a shutting-down replica finishes or cleanly abandons its lease instead of letting it silently expire, checkpointing granular enough to resume without redoing an hour of work but not so granular that the store becomes the bottleneck, and dashboards on delivery attempts and dead-letter arrivals so a crash loop shows up as a page instead of a mystery. None of that is a new idea, it’s the same lease-fencing-budget triangle we’ve been describing, just enforced by infrastructure instead of an in-process lock.
Host: So the in-memory store isn’t a toy you throw away for real deployment, it’s the spec you have to satisfy. That feels like the right place to leave it — the lab’s got 26 deterministic tests walking through fencing, redelivery, budget exhaustion, checkpoint resumption, every failure mode we talked about today, so if you want to see the guarantee actually hold under a crash instead of just believing it, that’s your starting point. Thanks for walking through it.
Not covered
The planner wanted these and found nothing in the source to support them:
- A live walkthrough of a Postgres- or Redis-backed production implementation replacing the in-memory store — the lab documents this as a required backend guarantee but doesn’t implement it
- A demonstration of multi-agent coordination using this task engine’s leases, as hinted at in the distributed-systems interview material
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Agent work is long-running, expensive, and frequently interrupted — a multi-step workflow that dies on step seven should not restart at step one, and a retried submission should not create a second task. This lab implements the delivery guarantees that make that true: leases with visibility timeouts, fencing tokens, idempotent submission, checkpointed resumption, and dead-lettering.
Source: labs/durable-agent-task-engine
Task lifecycle
Section titled “Task lifecycle”stateDiagram-v2
[*] --> Pending: submit (deduplicated by Idempotency-Key)
Pending --> Leased: lease() issues a fencing token
note right of Leased
Invisible to other workers only while
the lease is held. The visibility
timeout IS the crash-recovery
mechanism — no supervisor involved.
A stale worker's late ack, fail, or
checkpoint is rejected once its lease
has been reclaimed (lease fencing).
end note
Leased --> Succeeded: ack(valid token)
Leased --> Pending: fail(valid token), attempts remain
Leased --> Pending: lease expires (worker crashed)
Leased --> DeadLetter: delivery attempts exhausted
note right of DeadLetter
Counted by DELIVERY attempts, not
explicit failures — a worker that
crashes before it can call fail()
must still stop retrying eventually.
end note
DeadLetter --> Pending: operator requeue (never automatic)
Succeeded --> [*]What it demonstrates
Section titled “What it demonstrates”- The visibility timeout as crash recovery. A leased task is invisible to other workers only while its lease is held. If a worker dies, nobody renews the lease, it expires, and the task becomes available again — no heartbeat supervisor, no external watchdog, no liveness service.
- Lease fencing. A worker that resumes past its lease expiry is a zombie: it still believes it
owns work that has already been reissued. Its late
ack,fail, orcheckpointis rejected because its token is stale, which is what stops the task being completed twice. - Idempotent submission, separate from idempotent handling. An
Idempotency-Keydeduplicates task creation. It does nothing about the handler running twice, because at-least-once is the guarantee that is actually achievable — the two concerns are deliberately not conflated. - Dead-lettering by delivery attempt, not explicit failure. A task whose worker crashes before
it can call
fail()never records a failure. Counting deliveries instead of failures is what stops such a task retrying forever. - Checkpointed resumption. A handler calls
ctx.save_checkpoint({...})after each unit of progress, and the checkpoint is handed back throughctx.checkpointon the next attempt — whether that attempt came from a retry or from crash-driven redelivery. - Bounded concurrency with graceful drain.
shutdown()stops leasing new work and waits for in-flight tasks, so a deploy does not manufacture the exact crash the rest of the system exists to survive. - Deterministic tests for all of it. A clock fixture makes lease expiry and backoff instant, so 26 tests covering every failure mode above run in under a second with no real sleeps.
Why counting deliveries beats counting failures
Section titled “Why counting deliveries beats counting failures”This is the subtle one, and it is worth reading the two policies side by side.
A retry budget spent on explicit failures only decrements when a handler catches an error and
reports it. That works for a handler that raises — and fails completely for a handler whose process
is killed mid-task. The task’s lease expires, it returns to Pending with its budget untouched,
and it is redelivered forever. A poison task that reliably segfaults its worker becomes an infinite
loop that also takes down a worker each time around.
Counting delivery attempts closes that: the budget decrements when the task is handed out, so crash-driven redelivery is bounded by exactly the same limit as explicit failure. The task dead-letters, an operator sees it, and the loop terminates.
Run it
Section titled “Run it”cd labs/durable-agent-task-enginepython3.12 -m venv .venvsource .venv/bin/activatepip install -e '.[dev]'uvicorn task_queue.app:app --reloadThe app starts an in-process demo worker whose handler fails its first two attempts and succeeds on the third, so retries, backoff, and checkpointing are all observable from the API.
# Submit — the idempotency key is requiredcurl -s -X POST localhost:8000/v1/queues/orders/tasks \ -H 'content-type: application/json' \ -H 'Idempotency-Key: order-42-confirmation' \ -d '{"payload": {"order_id": "42"}, "max_attempts": 5}'
# Resubmitting the same key returns the original task, not a duplicatecurl -s http://localhost:8000/v1/tasks/<task-id>
# Dead-letter inspection and operator requeuecurl -s localhost:8000/v1/queues/orders/dead-lettercurl -s -X POST localhost:8000/v1/tasks/<task-id>/requeueVerify it
Section titled “Verify it”pytest # 26 testsruff check .mypy srcPrincipal-level discussion points
Section titled “Principal-level discussion points”- A visibility timeout implements crash recovery without a separate supervisor: if nobody proves they are still working, the work becomes available again. Fewer moving parts than a heartbeat service, and one less thing that can fail independently.
- Idempotent submission and idempotent handlers are different concerns. Deduplicating by key avoids duplicate tasks; the handler still has to be safe to run twice, because at-least-once is the achievable guarantee and exactly-once is not.
- Lease fencing is what makes at-least-once safe rather than merely tolerable — without it, a zombie worker’s late acknowledgement silently corrupts the state machine.
- Dead-lettering is an operator surface, not an automatic recovery path. A task usually lands there because a dependency is broken or a human needs to decide, not because it deserves another blind retry.
- This store’s single lock is correct for one process. A multi-replica deployment needs the same
atomicity from the backing store — row-level locking with
SKIP LOCKED, or an atomic Lua script — not from application code, which is why the in-memory implementation is written as a specification rather than a placeholder.
Related
Section titled “Related”- Architecture: Durable Agent Execution — the design-review companion to this lab: constraints, cost, observability, and deployment checklist.
- Module 2: Distributed Systems — idempotency, retries, and delivery guarantees this lab implements.
- Module 5: Agent Engineering — the long-running agent work that motivates checkpointed resumption.
labs/async-ai-gateway— theproduction-readyreference this lab is measured against.