LangGraph Checkpoint Cost
Read the transcript
1. Checkpointing is the whole bet
Host: So last time we built the agent loop by hand — a while loop, a step budget, a tool executor, the whole thing on a clipboard. Today we’re talking about LangGraph, which builds that same loop, but as an explicit graph instead of a loop you wrote yourself. And the question I want to open with is: why does that framing choice matter enough for a whole episode?
Guest: Because once control flow becomes data instead of a buried if-statement, you get something almost for free: after every single step, the graph checkpoints the full state to a thread ID. And that one mechanism is simultaneously your crash recovery and your human-in-the-loop approval gate — there’s no separate code path for ‘resume after the process died’ versus ‘resume after a human clicked approve.’ It’s the same load-the-last-checkpoint call either way.
Host: Which sounds like a clean win — one mechanism, two capabilities. But ‘persist the full state after every step’ is doing a lot of quiet work in that sentence, and that’s exactly what we’re going to spend this episode pulling apart: how that single checkpointing mechanism actually delivers both crash recovery and human-in-the-loop approval, and where the architecture of a graph as data makes that possible.
2. Building an instrument you can trust
Host: So before we get to the actual numbers, I want to talk about the measuring stick itself — because you built this lab’s instrument twice. What went wrong with the first version?
Guest: The first version just measured the bytes passed to the put function and called that the checkpoint cost. For most channel types that’s fine, because that put function is serializing the node’s actual output. But for the delta channel, its checkpoint method returns a sentinel value — missing — so the put function only ever sees that sentinel plus some remainder and metadata. The channel’s real per-step state never shows up there at all; it’s written separately through the put writes function into the pending-writes table. So the first instrument produced a clean, plausible table that was quietly missing the majority of what it claimed to total — and I had to write a test, one specifically checking that delta channel state is captured by put writes and not by put, to catch that blind spot and fail against it.
Host: Okay, so the fix is obvious then — just add the put bytes and the put writes bytes together and you get the true total. Why didn’t you do that?
Guest: Because that sum is only honest for one channel type. For every channel except the delta channel, the put writes function re-serializes the same node output that the put function already counted — so summing them double-counts the same bytes twice. The delta channel is the one case where the two columns are actually disjoint, so the sum means something there. A single ‘total bytes’ column would have to lie for one arm or the other, so the instrument keeps them as two separate columns and makes you reason about which one is doing the real work for each channel.
3. The headline: quadratic against linear
Host: Okay, so let’s put the actual numbers on the table. What do you get at 100 steps once you stop pre-summing and just look at each arm honestly?
Guest: The plain accumulating reducer’s put bytes hit about 1.34 megabytes at 100 steps. The DeltaChannel arm, doing the same accumulation, lands its honest total — put plus writes — at around 60 kilobytes. That’s over 22 times smaller for identical payloads.
Host: That’s a big constant-factor win, but you said earlier the shape matters more than the constant. What’s actually happening between 50 and 100 steps?
Guest: The accumulating channel’s put bytes go from 345,465 to 1,337,990 — that’s roughly quadrupling when the step count only doubles, because every write re-serializes the whole ever-growing list. DeltaChannel’s total goes from 30,251 to 59,951 — it just doubles, matching the step count. One is quadratic, the other is linear, and that’s the number that survives no matter how you slice the accounting.
4. The control arm, and its limits
Host: So you built a third arm, one that just replaces state instead of accumulating it. What’s that one for, if it’s not trying to win the comparison?
Guest: It’s the control. At every step count — 10, 25, 50, 100 — its final step bytes stay pinned at 538. Compare that to the accumulating arm hitting 26,181 bytes at 100 steps, and you can see the growth there is purely about accumulation, not some fixed overhead LangGraph charges per checkpoint.
Host: Okay, so naturally I want to ask: does DeltaChannel beat that flat baseline or not?
Guest: And that’s exactly the question this page refuses to answer, on purpose. Under put-only, put-plus-writes, or actually-retained-bytes, delta comes out lower — but under one asymmetric accounting it comes out 7% higher, and an earlier draft claimed it both ways before we caught that. A number that flips sign depending on which rule you pick isn’t a finding, it’s an artefact, so the control stays a baseline, not a contender.
5. Where the saved write cost goes: resume
Host: So DeltaChannel’s cheap writes aren’t free, they’re deferred — somebody pays on the way back in. Where does that bill actually land?
Guest: On resume. Reconstructing state means replaying ancestor writes through the reducer unless a full snapshot is close behind, and snapshot_frequency is the dial — the CLI sweeps 5, 25, 100, and 10000 at a 100-step run, that last one so large no snapshot ever gets written, so every resume replays the whole thing. And deliberately, this page publishes no table of those timings, because across twelve runs on the same build the resume figures weren’t stable — one draft quoted a single run as the number, another fit a two-decimal interval to six runs that the next five blew past. A figure that’s beaten two attempts to pin it doesn’t get a third shot at false precision. What survives is two claims: scale — every configuration resumes in tens to low hundreds of microseconds at 100 steps, well under a millisecond unloaded — and ordering — the never-snapshot sweep point was slowest in every single run, because it’s the deepest replay. That’s it, that’s what reproduces.
Host: And the number you’re actually timing when you measure that — it’s not pure replay, is it?
Guest: Right, it’s get_state latency, the whole call — checkpoint fetch, full deserialization, the ancestor walk, prepare_next_tasks, dict of subgraphs. Replay is one piece inside substantial fixed overhead, which is exactly why a huge difference in replay depth doesn’t translate into a proportionally huge difference in wall time. And watch for two traps if you rerun it yourself: the first measurement in a fresh process reads high from warm-up, and on a loaded machine one run put the never-snapshot row at 3.69 milliseconds — but every row inflated together that time, which is contention, not a statement about replay depth.
6. Beta, and what the reference page doesn’t say
Host: So if DeltaChannel fixes the quadratic write cost, why isn’t that just the answer — swap the channel, done? The lookup page calls it ‘the fix’ with no asterisk.
Guest: Because the docstring itself won’t say that. It says threads already written with it are ‘expected to remain readable’ — expected, not guaranteed — and that the on-disk representation and API around it may still change. We had to reach into undocumented internals just to measure it: a function for pulling delta channel history, an internal snapshot blob shape, and a counter tracking updates since the last snapshot — and those are named explicitly as not yet stable.
Host: So anyone reading delta state back out directly, not through get_state, is standing on ground that could move under them.
Guest: Exactly, and that’s before you get to the two conditions we haven’t even covered yet — it needs a batching-invariant reducer it can’t verify you’ve given it, and in the pinned release its write path can deadlock. The effect is real, the write-path numbers back it up. It’s just not the unconditional fix the page implies.
7. A real deadlock, and the wrong explanation for it
Host: Okay, deadlock. Walk me through it, because ‘shared threadpool’ is the explanation everyone’s going to reach for first — node execution and checkpoint writes sitting in one bounded executor.
Guest: Right, and that explanation is wrong, which is why it’s worth stating precisely. Yes, pregel’s runner submits both node work and checkpoint puts into the same bounded ThreadPoolExecutor — but that alone doesn’t deadlock anything. Each step’s put waits on the previous step’s put future, and in a strictly-FIFO pool the earliest task in that chain has nothing left to wait on, so it gets a worker, unblocks, and the chain drains fine.
Host: So if it’s not the shared pool, what actually traps the workers?
Guest: It’s an order inversion specific to DeltaChannel. Before waiting on its predecessor, the put-after-previous call drains a list of delta write futures and waits on whatever’s in that list at execution time, not at submission time — the loop keeps advancing while earlier puts sit queued, so a task can end up waiting on futures that were submitted to the pool after it was. Enough of those parked at once and every worker is waiting on a future that can’t start — that’s circular wait. We ran an unpatched 100-step sweep to check it wasn’t theoretical, it hung, and we killed it after five minutes; the one-line max_concurrency workaround on that arm is load-bearing, not a precaution, and I checked the issue tracker — nothing matching it turned up.
8. The failure a dashboard would never show
Host: Okay, the deadlock was loud — it hangs, you notice, you kill it. What’s the failure that wouldn’t announce itself at all?
Guest: A non-associative reducer. On replay, DeltaChannel folds writes in bigger batches than they were originally produced in, and if your reducer cares about batching at all, you get a different answer with zero error raised. We measured it on a twelve-step run — twenty-five items live at the time, thirteen items after replay. Nothing crashes, nothing logs, you just silently have the wrong state and no mechanism tells you.
Host: So that’s the quiet one. Now, you also drew a line earlier in the writeup between two claims that sound almost identical — walk me through that distinction before we close this out.
Guest: Right — we proved a fresh, independently compiled graph can read back *finished* state through get_state alone, no Command, no interrupt call, its node functions never even ran. That’s real and it’s clean. What we did not prove is the bigger claim people actually want: that a still-pending interrupted checkpoint can be picked up by a fresh process and driven to completion with no Command. Finished state surviving into a new object is not the same experiment as resuming a pending one, and only the first one is in our suite.
9. What a principal engineer takes from this
Host: So let’s pull up from the specifics — if you’re a principal engineer skimming this whole investigation, what are the load-bearing takeaways, the things you’d actually say in a design review?
Guest: Four things. First, the instrument is part of the result — our own first table was arithmetically correct over an incomplete number, and it still read as a finding, so always ask what call sites a cost meter wraps and what it never sees. Second, the fix is a state-model rewrite, and the beta risk is real but narrower than it sounds — LangGraph says the API and the on-disk representation may change, and that written threads are merely expected, not guaranteed, to stay readable, which is a smaller risk than it feels like until you’re the one reading delta state directly. Third, and this is the one that should worry you most, a silent correctness failure beats a loud cost failure every time — the quadratic curve shows up on a dashboard, but a non-associative reducer quietly producing thirteen items where twenty-five lived shows up nowhere. And fourth, deferred cost doesn’t vanish, it changes owners — write cost is paid by the service doing the work, but resume cost is paid by whoever’s blocked on get state, often a user, often exactly on the recovery path, which is the worst possible moment to find out.
Host: That last one lands hardest for me — you don’t get to choose when the bill comes due, you just get to choose who’s holding it. Which is really the whole shape of this lab: the cost was always honest, it just wasn’t always where anyone was looking.
10. Go run it yourself
Host: So if someone wants to stop taking our word for it, where do they go? Walk me through actually running this thing.
Guest: It’s the langgraph-checkpoint-cost lab — set up a venv, install with the dev extras, and run ruff, mypy, and pytest before the CLI itself. Nineteen tests, ruff clean, and mypy strict runs over the tests directory too, not just source, because if you’re going to trust a number you have to trust the code that produced it. Then the CLI prints the byte tables and the resume sweep live on your machine, no published numbers to just believe. And if you want the bigger picture, Module 7 covers the LangGraph claims we measured, the LangGraph lookup is the DeltaChannel recommendation this lab qualifies, and the Durable Agent Task Engine lab shows what checkpointed resumption looks like as a full system rather than one channel.
Host: Perfect place to send people. That’s the lab, that’s the episode — go run it, look at your own tables, and decide for yourself where your bill is actually landing. Thanks for walking through all of it with me.
Not covered
The planner wanted these and found nothing in the source to support them:
- Checkpoint cost benchmarks against a production Postgres-backed checkpointer rather than the lab’s local saver
- Cost comparison of DeltaChannel across different LangGraph minor versions beyond the pinned 1.2.11
- Guidance on migrating an existing production thread history from a plain accumulating channel to DeltaChannel
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Module 7 says a checkpointer persists state every step, and the LangGraph lookup says large accumulated state gets slower every step because the whole value is re-serialized per checkpoint. Both are claims about cost. This lab puts bytes on them, on the actual serialization path LangGraph uses, and then measures what the recommended fix costs on the other side.
Source: labs/langgraph-checkpoint-cost
The finding
Section titled “The finding”Three graphs, identical payloads, swept over step count. put bytes is what the checkpointer
serializes through put(); writes bytes is what it serializes through put_writes() into the
pending-writes table. They are reported separately and never pre-summed, for a reason the next
section explains.
Accumulating — a plain reducer appending to a list in state:
steps put bytes writes bytes put+writes final step bytes items 10 17655 2612 20267 2869 10/10 25 92015 6527 98542 6756 25/25 50 345465 13052 358517 13231 50/50 100 1337990 26102 1364092 26181 100/100The same accumulation through DeltaChannel:
steps put bytes writes bytes put+writes final step bytes items 10 3879 2612 6491 318 10/10 25 8874 6527 15401 318 25/25 50 17199 13052 30251 318 50/50 100 33849 26102 59951 318 100/100At 100 steps: roughly 1.34 MB against roughly 60 KB — over 22× smaller. And the shape differs,
not just the constant. Between 50 and 100 steps the accumulating channel’s put bytes roughly
quadruples (345,465 → 1,337,990), which is what re-serializing an ever-growing list on every write
costs. DeltaChannel’s honest total merely doubles (30,251 → 59,951). Quadratic against linear.
This is the result to carry away, and it is the one that survives every accounting we tried.
The control, and what it is not for
Section titled “The control, and what it is not for”A third arm replaces state each step instead of accumulating it, so its per-step cost is flat by construction:
steps put bytes writes bytes put+writes final step bytes items 10 6000 2612 8612 538 1/1 25 14295 6527 20822 538 1/1 50 28120 13052 41172 538 1/1 100 55770 26102 81872 538 1/1Note the final step bytes column: 538 at every step count, against the accumulating arm’s 26,181
at 100 steps. That is the control doing its job — showing what a checkpoint costs when nothing
accumulates, so the growth in the other two arms is attributable to accumulation rather than to
LangGraph’s fixed per-checkpoint overhead.
This page draws no verdict comparing DeltaChannel to that control, in either direction. Under
put-only, put-plus-writes, and actually-retained-bytes accountings (the last measured separately —
the CLI prints only the first two, not the third), delta comes out lower; under one asymmetric
accounting it comes out 7% higher. A 7% gap that flips sign with the accounting
rule is not a finding, it is an artefact of the rule, and an earlier draft of this lab’s report
claimed it in each direction before that became clear. The control is here as a flat-cost baseline.
It is not a contender.
Why put and writes are counted separately
Section titled “Why put and writes are counted separately”DeltaChannel.checkpoint() returns MISSING. Its committed state never appears in a checkpoint’s
channel_values, so put() sees only a sentinel, the checkpoint remainder, and metadata — the
channel’s actual per-step state lives in the pending-writes table, written through put_writes().
The first version of this lab’s instrument measured put() alone. It produced a clean, plausible,
wrong table, because it was totalling a number that was missing the majority of what it claimed
to total. test_delta_channel_state_is_captured_by_put_writes_not_put exists specifically to fail
against that blind instrument.
The two columns are still not summed inside the measuring saver. For every channel here except
DeltaChannel, put_writes re-serializes the same step’s node output that put() already counted,
so adding them double-counts. Only for DeltaChannel are the two disjoint and only there does the
sum mean anything. A single “total bytes” column would have to be wrong for one arm or the other.
The other side of the trade: resume
Section titled “The other side of the trade: resume”A channel that makes writes cheap by deferring work does that work somewhere. For DeltaChannel
it is on read — reconstructing state means replaying ancestor writes through the reducer, unless a
full snapshot is close enough behind. snapshot_frequency is the dial, and the CLI sweeps it at
100 steps across 5, 25, 100, and 10000 — the last being large enough that, over 100 steps, no
snapshot is written at all, so every resume replays the full run.
This page deliberately publishes no table of those timings, and that absence is the result. The three byte tables above are bit-for-bit identical across all twelve runs on record, all on the same installed build — that sameness is a property of a deterministic serialization path, not evidence of reproduction across independently sourced environments, so it is stated as a run count and no more. The resume timings are not identical, and two earlier drafts of this page were wrong about them in opposite directions — first quoting one run’s figures as if they were the number, then fitting a two-decimal interval to six runs that the next five runs fell outside. A figure that has defeated two attempts to pin it should not be printed to two decimal places a third time. What twelve runs do support is two claims, and the section rests on them:
- Scale. Every configuration resumes in tens to low hundreds of microseconds at 100 steps —
well under a millisecond in every run on an unloaded machine (see the contention note below). That
is the honest size of
DeltaChannel’s deferred cost: real, measurable, and small at this scale. - Ordering. The never-snapshot sweep point was the slowest in every run, because it is the deepest replay. That is what the section actually argues, and it is the part that reproduces.
Run the CLI yourself and you get the raw figures, which is the right place for them.
What you time is get_state latency, not replay time. ResumeResult.reconstruct_seconds times
the entire get_state call — checkpoint fetch, full deserialization, channels_from_checkpoint
(the ancestor walk this lab cares about), prepare_next_tasks, and dict(get_subgraphs()).
Ancestor replay is one component of a call with substantial fixed overhead, which is exactly why a
large difference in replay depth does not produce a proportionally large difference in wall time.
Two things a reader reproducing this will hit. The first timing in a fresh process runs high — it is the first measurement taken after import, and warm-up lands on it, so an initial row can read several times its steady-state value without anything being wrong. And one run on a loaded machine put the never-snapshot row at 3.69 ms; in that run every row inflated together, which is the signature of contention rather than of anything about replay depth.
DeltaChannel is beta, and the Reference page oversells it
Section titled “DeltaChannel is beta, and the Reference page oversells it”The LangGraph lookup currently calls DeltaChannel “the fix for
large-state runs that get slower as they progress,” with no qualification. The write-path
measurement above supports the effect. It does not support the unqualified recommendation, and
this lab’s contact with the API says why:
- The channel’s own docstring states that its API and its on-disk representation may change. It
is worth quoting what that does and does not promise: threads already written with
DeltaChannel“are expected to remain readable”, and the instability is scoped to the surrounding contract. Note the hedge — expected, not guaranteed. It is a stated expectation about existing threads, not a compatibility guarantee, and the representation underneath them is explicitly not stable. - That surrounding contract is what this lab had to reach for to measure the channel at all —
get_delta_channel_history, the_DeltaSnapshotblob shape,counters_since_delta_snapshot— and it is named in the docstring as not yet stable. Anything reading delta state back out itself, rather than throughget_state, is coupled to internals that may move. - It requires a batching-invariant reducer, and cannot check that you gave it one. See below.
- In the pinned release, its write path can deadlock. See below.
None of that makes it the wrong choice, and the stated expectation above makes it a smaller bet than “beta storage format” usually implies. It makes it a choice with conditions attached, which “the fix” does not communicate at all.
A real deadlock in langgraph 1.2.11
Section titled “A real deadlock in langgraph 1.2.11”The DeltaChannel write path in langgraph==1.2.11 can deadlock, and the mechanism is worth
stating precisely because the obvious explanation is wrong.
Node execution and checkpoint writes share one bounded ThreadPoolExecutor
(pregel/_runner.py:260 submits node work into the same pool). That alone does not deadlock
anything: each step’s checkpoint put waits on the previous step’s put future, and in a
strictly-FIFO pool the earliest task in that chain has nothing left to wait on, gets a worker, and
unblocks the next. A reader who knows how ThreadPoolExecutor dispatches would correctly reject
“both waits happen in the same bounded pool” as an explanation.
The actual condition is an order inversion, and it is specific to DeltaChannel. Before waiting
on its predecessor, _checkpointer_put_after_previous drains self._delta_write_futs and waits on
whatever is in that list at execution time, not at submission time
(pregel/_loop.py:1538-1540). The main loop keeps advancing while earlier put tasks sit queued, so
a task that finally gets a worker can find itself waiting on put_writes futures that were
submitted to the pool after it was — occupying a worker while blocked on something queued behind
it. With enough of those parked at once, every worker is waiting on a future that cannot start.
That is circular wait, not slow drain.
This is not theoretical. An unpatched 100-step delta sweep was run, hung, and was killed after
five minutes. The one-line max_concurrency on the delta arm’s invoke config is a load-bearing
workaround, not a precaution — and byte counts under it match unpatched runs at step counts where
the unpatched run does finish, so the workaround does not perturb the measurement.
The LangGraph issue tracker was searched for this and nothing matching was found — first on 2026-08-30, again on 2026-10-08. That is a statement about the search, not about the bug’s reporting status.
Still present in 1.2.14
Section titled “Still present in 1.2.14”Re-checked on 2026-10-08 against langgraph==1.2.14, three patch releases later. 1.2.13 shipped
five DeltaChannel fixes, none of them this one: _checkpointer_put_after_previous is
byte-for-byte unchanged (now pregel/_loop.py:1604-1606). The lab’s own delta graph, run at 60
steps with no max_concurrency, hung three times out of three; with the workaround it completed
thirty times out of thirty. The lab’s full test suite passes on 1.2.14 unchanged. Its pin
stays at 1.2.11, because that is the release its byte counts were measured on.
What else the tests pin
Section titled “What else the tests pin”A non-associative reducer silently reconstructs different state. Replay folds writes in larger
batches than they were produced in, so a reducer sensitive to batching returns something else after
a resume. Measured on a 12-step run: 25 items live, 13 items after replay, no error raised
anywhere. DeltaChannel requires a batching-invariant reducer and has no way to check that you
supplied one — the failure is silent, and it is silently wrong state, not a crash.
Human approval resumes through interrupt() and Command(resume=...). The graph pauses, the
run returns with __interrupt__ set, and the original compiled graph resumes it to completion.
Finished state survives into a completely fresh graph object. A second graph, compiled
independently against the same saver and sharing nothing in-process with the first, reads the
completed values back through get_state alone — no Command, no interrupt(), its node functions
never executed.
Those are two pinned facts, and this page deliberately stops short of Module 7’s fuller claim.
What the suite does not prove is that a still-pending interrupted checkpoint can be driven to
completion by a fresh process with no Command. The crash-recovery arm reads finished state
from a fresh object; it does not resume a pending one. That is a narrower result than
“human-in-the-loop and crash recovery are one mechanism,” and the narrower result is the one the
tests support.
Run it
Section titled “Run it”cd labs/langgraph-checkpoint-costuv venv .venv && uv pip install --python .venv/bin/python -e '.[dev]'./.venv/bin/python -m ruff check . && ./.venv/bin/python -m mypy src tests && ./.venv/bin/python -m pytest -q./.venv/bin/python -m checkpoint_cost.cli19 tests, ruff clean, mypy --strict clean. Unlike this repo’s other labs, mypy runs over
tests as well as src: a chunk of this lab’s value is “trust these numbers,” so the code
producing them is held to the same bar as the code under test. The last command prints the three
byte tables above plus the resume sweep, whose raw timings are deliberately not published here.
Principal-level discussion points
Section titled “Principal-level discussion points”- The instrument is part of the result. This lab’s first table was arithmetically correct over an incomplete number, and it read as a finding. Ask of any cost measurement: what call sites does the meter wrap, and what does the system do that it never sees?
- “The fix” is a state-model rewrite, and the beta risk is narrower than it sounds. Adopting
DeltaChannellate means changing how state is modelled, under load, on a running system. The storage risk is the smaller half, but it is not zero: LangGraph says the API and the on-disk representation may change, and that threads already written are expected — not guaranteed — to remain readable. What it names outright as unstable is the surrounding contract, which bites hardest if you read delta state yourself, and the representation underneathget_stateis itself only expected, not guaranteed, to stay readable. Know which of the two risks you are actually taking, and how firm the assurance under it is, before you quote either to a review. - A silent correctness failure beats a loud cost failure, every time. The quadratic write curve is visible on a dashboard. A non-associative reducer producing 13 items where 25 lived is visible to nobody.
- Deferred cost is still cost, and it moves to a different owner. Write cost is paid by the
service doing the work; resume cost is paid by whoever is waiting on
get_state— often a user, often on the recovery path, which is the worst moment to discover it.
Related
Section titled “Related”- Module 7: LangGraph — state, reducers, and the checkpointing claims this lab measures.
- LangGraph — the lookup whose
DeltaChannelrecommendation this lab qualifies. - Durable Agent Task Engine — checkpointed resumption and retry semantics as a system rather than as a channel.