Skip to content

LangGraph Checkpoint Cost

Listen to this page14:28
Read the transcript

1. Checkpointing is the whole bet

Host: So last time we built the agent loop by hand — a while loop, a step budget, a tool executor, the whole thing on a clipboard. Today we’re talking about LangGraph, which builds that same loop, but as an explicit graph instead of a loop you wrote yourself. And the question I want to open with is: why does that framing choice matter enough for a whole episode?

Guest: Because once control flow becomes data instead of a buried if-statement, you get something almost for free: after every single step, the graph checkpoints the full state to a thread ID. And that one mechanism is simultaneously your crash recovery and your human-in-the-loop approval gate — there’s no separate code path for ‘resume after the process died’ versus ‘resume after a human clicked approve.’ It’s the same load-the-last-checkpoint call either way.

Host: Which sounds like a clean win — one mechanism, two capabilities. But ‘persist the full state after every step’ is doing a lot of quiet work in that sentence, and that’s exactly what we’re going to spend this episode pulling apart: how that single checkpointing mechanism actually delivers both crash recovery and human-in-the-loop approval, and where the architecture of a graph as data makes that possible.

2. Building an instrument you can trust

Host: So before we get to the actual numbers, I want to talk about the measuring stick itself — because you built this lab’s instrument twice. What went wrong with the first version?

Guest: The first version just measured the bytes passed to the put function and called that the checkpoint cost. For most channel types that’s fine, because that put function is serializing the node’s actual output. But for the delta channel, its checkpoint method returns a sentinel value — missing — so the put function only ever sees that sentinel plus some remainder and metadata. The channel’s real per-step state never shows up there at all; it’s written separately through the put writes function into the pending-writes table. So the first instrument produced a clean, plausible table that was quietly missing the majority of what it claimed to total — and I had to write a test, one specifically checking that delta channel state is captured by put writes and not by put, to catch that blind spot and fail against it.

Host: Okay, so the fix is obvious then — just add the put bytes and the put writes bytes together and you get the true total. Why didn’t you do that?

Guest: Because that sum is only honest for one channel type. For every channel except the delta channel, the put writes function re-serializes the same node output that the put function already counted — so summing them double-counts the same bytes twice. The delta channel is the one case where the two columns are actually disjoint, so the sum means something there. A single ‘total bytes’ column would have to lie for one arm or the other, so the instrument keeps them as two separate columns and makes you reason about which one is doing the real work for each channel.

3. The headline: quadratic against linear

Host: Okay, so let’s put the actual numbers on the table. What do you get at 100 steps once you stop pre-summing and just look at each arm honestly?

Guest: The plain accumulating reducer’s put bytes hit about 1.34 megabytes at 100 steps. The DeltaChannel arm, doing the same accumulation, lands its honest total — put plus writes — at around 60 kilobytes. That’s over 22 times smaller for identical payloads.

Host: That’s a big constant-factor win, but you said earlier the shape matters more than the constant. What’s actually happening between 50 and 100 steps?

Guest: The accumulating channel’s put bytes go from 345,465 to 1,337,990 — that’s roughly quadrupling when the step count only doubles, because every write re-serializes the whole ever-growing list. DeltaChannel’s total goes from 30,251 to 59,951 — it just doubles, matching the step count. One is quadratic, the other is linear, and that’s the number that survives no matter how you slice the accounting.

4. The control arm, and its limits

Host: So you built a third arm, one that just replaces state instead of accumulating it. What’s that one for, if it’s not trying to win the comparison?

Guest: It’s the control. At every step count — 10, 25, 50, 100 — its final step bytes stay pinned at 538. Compare that to the accumulating arm hitting 26,181 bytes at 100 steps, and you can see the growth there is purely about accumulation, not some fixed overhead LangGraph charges per checkpoint.

Host: Okay, so naturally I want to ask: does DeltaChannel beat that flat baseline or not?

Guest: And that’s exactly the question this page refuses to answer, on purpose. Under put-only, put-plus-writes, or actually-retained-bytes, delta comes out lower — but under one asymmetric accounting it comes out 7% higher, and an earlier draft claimed it both ways before we caught that. A number that flips sign depending on which rule you pick isn’t a finding, it’s an artefact, so the control stays a baseline, not a contender.

5. Where the saved write cost goes: resume

Host: So DeltaChannel’s cheap writes aren’t free, they’re deferred — somebody pays on the way back in. Where does that bill actually land?

Guest: On resume. Reconstructing state means replaying ancestor writes through the reducer unless a full snapshot is close behind, and snapshot_frequency is the dial — the CLI sweeps 5, 25, 100, and 10000 at a 100-step run, that last one so large no snapshot ever gets written, so every resume replays the whole thing. And deliberately, this page publishes no table of those timings, because across twelve runs on the same build the resume figures weren’t stable — one draft quoted a single run as the number, another fit a two-decimal interval to six runs that the next five blew past. A figure that’s beaten two attempts to pin it doesn’t get a third shot at false precision. What survives is two claims: scale — every configuration resumes in tens to low hundreds of microseconds at 100 steps, well under a millisecond unloaded — and ordering — the never-snapshot sweep point was slowest in every single run, because it’s the deepest replay. That’s it, that’s what reproduces.

Host: And the number you’re actually timing when you measure that — it’s not pure replay, is it?

Guest: Right, it’s get_state latency, the whole call — checkpoint fetch, full deserialization, the ancestor walk, prepare_next_tasks, dict of subgraphs. Replay is one piece inside substantial fixed overhead, which is exactly why a huge difference in replay depth doesn’t translate into a proportionally huge difference in wall time. And watch for two traps if you rerun it yourself: the first measurement in a fresh process reads high from warm-up, and on a loaded machine one run put the never-snapshot row at 3.69 milliseconds — but every row inflated together that time, which is contention, not a statement about replay depth.

6. Beta, and what the reference page doesn’t say

Host: So if DeltaChannel fixes the quadratic write cost, why isn’t that just the answer — swap the channel, done? The lookup page calls it ‘the fix’ with no asterisk.

Guest: Because the docstring itself won’t say that. It says threads already written with it are ‘expected to remain readable’ — expected, not guaranteed — and that the on-disk representation and API around it may still change. We had to reach into undocumented internals just to measure it: a function for pulling delta channel history, an internal snapshot blob shape, and a counter tracking updates since the last snapshot — and those are named explicitly as not yet stable.

Host: So anyone reading delta state back out directly, not through get_state, is standing on ground that could move under them.

Guest: Exactly, and that’s before you get to the two conditions we haven’t even covered yet — it needs a batching-invariant reducer it can’t verify you’ve given it, and in the pinned release its write path can deadlock. The effect is real, the write-path numbers back it up. It’s just not the unconditional fix the page implies.

7. A real deadlock, and the wrong explanation for it

Host: Okay, deadlock. Walk me through it, because ‘shared threadpool’ is the explanation everyone’s going to reach for first — node execution and checkpoint writes sitting in one bounded executor.

Guest: Right, and that explanation is wrong, which is why it’s worth stating precisely. Yes, pregel’s runner submits both node work and checkpoint puts into the same bounded ThreadPoolExecutor — but that alone doesn’t deadlock anything. Each step’s put waits on the previous step’s put future, and in a strictly-FIFO pool the earliest task in that chain has nothing left to wait on, so it gets a worker, unblocks, and the chain drains fine.

Host: So if it’s not the shared pool, what actually traps the workers?

Guest: It’s an order inversion specific to DeltaChannel. Before waiting on its predecessor, the put-after-previous call drains a list of delta write futures and waits on whatever’s in that list at execution time, not at submission time — the loop keeps advancing while earlier puts sit queued, so a task can end up waiting on futures that were submitted to the pool after it was. Enough of those parked at once and every worker is waiting on a future that can’t start — that’s circular wait. We ran an unpatched 100-step sweep to check it wasn’t theoretical, it hung, and we killed it after five minutes; the one-line max_concurrency workaround on that arm is load-bearing, not a precaution, and I checked the issue tracker — nothing matching it turned up.

8. The failure a dashboard would never show

Host: Okay, the deadlock was loud — it hangs, you notice, you kill it. What’s the failure that wouldn’t announce itself at all?

Guest: A non-associative reducer. On replay, DeltaChannel folds writes in bigger batches than they were originally produced in, and if your reducer cares about batching at all, you get a different answer with zero error raised. We measured it on a twelve-step run — twenty-five items live at the time, thirteen items after replay. Nothing crashes, nothing logs, you just silently have the wrong state and no mechanism tells you.

Host: So that’s the quiet one. Now, you also drew a line earlier in the writeup between two claims that sound almost identical — walk me through that distinction before we close this out.

Guest: Right — we proved a fresh, independently compiled graph can read back *finished* state through get_state alone, no Command, no interrupt call, its node functions never even ran. That’s real and it’s clean. What we did not prove is the bigger claim people actually want: that a still-pending interrupted checkpoint can be picked up by a fresh process and driven to completion with no Command. Finished state surviving into a new object is not the same experiment as resuming a pending one, and only the first one is in our suite.

9. What a principal engineer takes from this

Host: So let’s pull up from the specifics — if you’re a principal engineer skimming this whole investigation, what are the load-bearing takeaways, the things you’d actually say in a design review?

Guest: Four things. First, the instrument is part of the result — our own first table was arithmetically correct over an incomplete number, and it still read as a finding, so always ask what call sites a cost meter wraps and what it never sees. Second, the fix is a state-model rewrite, and the beta risk is real but narrower than it sounds — LangGraph says the API and the on-disk representation may change, and that written threads are merely expected, not guaranteed, to stay readable, which is a smaller risk than it feels like until you’re the one reading delta state directly. Third, and this is the one that should worry you most, a silent correctness failure beats a loud cost failure every time — the quadratic curve shows up on a dashboard, but a non-associative reducer quietly producing thirteen items where twenty-five lived shows up nowhere. And fourth, deferred cost doesn’t vanish, it changes owners — write cost is paid by the service doing the work, but resume cost is paid by whoever’s blocked on get state, often a user, often exactly on the recovery path, which is the worst possible moment to find out.

Host: That last one lands hardest for me — you don’t get to choose when the bill comes due, you just get to choose who’s holding it. Which is really the whole shape of this lab: the cost was always honest, it just wasn’t always where anyone was looking.

10. Go run it yourself

Host: So if someone wants to stop taking our word for it, where do they go? Walk me through actually running this thing.

Guest: It’s the langgraph-checkpoint-cost lab — set up a venv, install with the dev extras, and run ruff, mypy, and pytest before the CLI itself. Nineteen tests, ruff clean, and mypy strict runs over the tests directory too, not just source, because if you’re going to trust a number you have to trust the code that produced it. Then the CLI prints the byte tables and the resume sweep live on your machine, no published numbers to just believe. And if you want the bigger picture, Module 7 covers the LangGraph claims we measured, the LangGraph lookup is the DeltaChannel recommendation this lab qualifies, and the Durable Agent Task Engine lab shows what checkpointed resumption looks like as a full system rather than one channel.

Host: Perfect place to send people. That’s the lab, that’s the episode — go run it, look at your own tables, and decide for yourself where your bill is actually landing. Thanks for walking through all of it with me.

Not covered

The planner wanted these and found nothing in the source to support them:

  • Checkpoint cost benchmarks against a production Postgres-backed checkpointer rather than the lab’s local saver
  • Cost comparison of DeltaChannel across different LangGraph minor versions beyond the pinned 1.2.11
  • Guidance on migrating an existing production thread history from a plain accumulating channel to DeltaChannel

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

Module 7 says a checkpointer persists state every step, and the LangGraph lookup says large accumulated state gets slower every step because the whole value is re-serialized per checkpoint. Both are claims about cost. This lab puts bytes on them, on the actual serialization path LangGraph uses, and then measures what the recommended fix costs on the other side.

Source: labs/langgraph-checkpoint-cost

Three graphs, identical payloads, swept over step count. put bytes is what the checkpointer serializes through put(); writes bytes is what it serializes through put_writes() into the pending-writes table. They are reported separately and never pre-summed, for a reason the next section explains.

Accumulating — a plain reducer appending to a list in state:

steps put bytes writes bytes put+writes final step bytes items
10 17655 2612 20267 2869 10/10
25 92015 6527 98542 6756 25/25
50 345465 13052 358517 13231 50/50
100 1337990 26102 1364092 26181 100/100

The same accumulation through DeltaChannel:

steps put bytes writes bytes put+writes final step bytes items
10 3879 2612 6491 318 10/10
25 8874 6527 15401 318 25/25
50 17199 13052 30251 318 50/50
100 33849 26102 59951 318 100/100

At 100 steps: roughly 1.34 MB against roughly 60 KB — over 22× smaller. And the shape differs, not just the constant. Between 50 and 100 steps the accumulating channel’s put bytes roughly quadruples (345,465 → 1,337,990), which is what re-serializing an ever-growing list on every write costs. DeltaChannel’s honest total merely doubles (30,251 → 59,951). Quadratic against linear.

This is the result to carry away, and it is the one that survives every accounting we tried.

A third arm replaces state each step instead of accumulating it, so its per-step cost is flat by construction:

steps put bytes writes bytes put+writes final step bytes items
10 6000 2612 8612 538 1/1
25 14295 6527 20822 538 1/1
50 28120 13052 41172 538 1/1
100 55770 26102 81872 538 1/1

Note the final step bytes column: 538 at every step count, against the accumulating arm’s 26,181 at 100 steps. That is the control doing its job — showing what a checkpoint costs when nothing accumulates, so the growth in the other two arms is attributable to accumulation rather than to LangGraph’s fixed per-checkpoint overhead.

This page draws no verdict comparing DeltaChannel to that control, in either direction. Under put-only, put-plus-writes, and actually-retained-bytes accountings (the last measured separately — the CLI prints only the first two, not the third), delta comes out lower; under one asymmetric accounting it comes out 7% higher. A 7% gap that flips sign with the accounting rule is not a finding, it is an artefact of the rule, and an earlier draft of this lab’s report claimed it in each direction before that became clear. The control is here as a flat-cost baseline. It is not a contender.

DeltaChannel.checkpoint() returns MISSING. Its committed state never appears in a checkpoint’s channel_values, so put() sees only a sentinel, the checkpoint remainder, and metadata — the channel’s actual per-step state lives in the pending-writes table, written through put_writes().

The first version of this lab’s instrument measured put() alone. It produced a clean, plausible, wrong table, because it was totalling a number that was missing the majority of what it claimed to total. test_delta_channel_state_is_captured_by_put_writes_not_put exists specifically to fail against that blind instrument.

The two columns are still not summed inside the measuring saver. For every channel here except DeltaChannel, put_writes re-serializes the same step’s node output that put() already counted, so adding them double-counts. Only for DeltaChannel are the two disjoint and only there does the sum mean anything. A single “total bytes” column would have to be wrong for one arm or the other.

A channel that makes writes cheap by deferring work does that work somewhere. For DeltaChannel it is on read — reconstructing state means replaying ancestor writes through the reducer, unless a full snapshot is close enough behind. snapshot_frequency is the dial, and the CLI sweeps it at 100 steps across 5, 25, 100, and 10000 — the last being large enough that, over 100 steps, no snapshot is written at all, so every resume replays the full run.

This page deliberately publishes no table of those timings, and that absence is the result. The three byte tables above are bit-for-bit identical across all twelve runs on record, all on the same installed build — that sameness is a property of a deterministic serialization path, not evidence of reproduction across independently sourced environments, so it is stated as a run count and no more. The resume timings are not identical, and two earlier drafts of this page were wrong about them in opposite directions — first quoting one run’s figures as if they were the number, then fitting a two-decimal interval to six runs that the next five runs fell outside. A figure that has defeated two attempts to pin it should not be printed to two decimal places a third time. What twelve runs do support is two claims, and the section rests on them:

  • Scale. Every configuration resumes in tens to low hundreds of microseconds at 100 steps — well under a millisecond in every run on an unloaded machine (see the contention note below). That is the honest size of DeltaChannel’s deferred cost: real, measurable, and small at this scale.
  • Ordering. The never-snapshot sweep point was the slowest in every run, because it is the deepest replay. That is what the section actually argues, and it is the part that reproduces.

Run the CLI yourself and you get the raw figures, which is the right place for them.

What you time is get_state latency, not replay time. ResumeResult.reconstruct_seconds times the entire get_state call — checkpoint fetch, full deserialization, channels_from_checkpoint (the ancestor walk this lab cares about), prepare_next_tasks, and dict(get_subgraphs()). Ancestor replay is one component of a call with substantial fixed overhead, which is exactly why a large difference in replay depth does not produce a proportionally large difference in wall time.

Two things a reader reproducing this will hit. The first timing in a fresh process runs high — it is the first measurement taken after import, and warm-up lands on it, so an initial row can read several times its steady-state value without anything being wrong. And one run on a loaded machine put the never-snapshot row at 3.69 ms; in that run every row inflated together, which is the signature of contention rather than of anything about replay depth.

DeltaChannel is beta, and the Reference page oversells it

Section titled “DeltaChannel is beta, and the Reference page oversells it”

The LangGraph lookup currently calls DeltaChannel “the fix for large-state runs that get slower as they progress,” with no qualification. The write-path measurement above supports the effect. It does not support the unqualified recommendation, and this lab’s contact with the API says why:

  • The channel’s own docstring states that its API and its on-disk representation may change. It is worth quoting what that does and does not promise: threads already written with DeltaChannel “are expected to remain readable”, and the instability is scoped to the surrounding contract. Note the hedge — expected, not guaranteed. It is a stated expectation about existing threads, not a compatibility guarantee, and the representation underneath them is explicitly not stable.
  • That surrounding contract is what this lab had to reach for to measure the channel at all — get_delta_channel_history, the _DeltaSnapshot blob shape, counters_since_delta_snapshot — and it is named in the docstring as not yet stable. Anything reading delta state back out itself, rather than through get_state, is coupled to internals that may move.
  • It requires a batching-invariant reducer, and cannot check that you gave it one. See below.
  • In the pinned release, its write path can deadlock. See below.

None of that makes it the wrong choice, and the stated expectation above makes it a smaller bet than “beta storage format” usually implies. It makes it a choice with conditions attached, which “the fix” does not communicate at all.

The DeltaChannel write path in langgraph==1.2.11 can deadlock, and the mechanism is worth stating precisely because the obvious explanation is wrong.

Node execution and checkpoint writes share one bounded ThreadPoolExecutor (pregel/_runner.py:260 submits node work into the same pool). That alone does not deadlock anything: each step’s checkpoint put waits on the previous step’s put future, and in a strictly-FIFO pool the earliest task in that chain has nothing left to wait on, gets a worker, and unblocks the next. A reader who knows how ThreadPoolExecutor dispatches would correctly reject “both waits happen in the same bounded pool” as an explanation.

The actual condition is an order inversion, and it is specific to DeltaChannel. Before waiting on its predecessor, _checkpointer_put_after_previous drains self._delta_write_futs and waits on whatever is in that list at execution time, not at submission time (pregel/_loop.py:1538-1540). The main loop keeps advancing while earlier put tasks sit queued, so a task that finally gets a worker can find itself waiting on put_writes futures that were submitted to the pool after it was — occupying a worker while blocked on something queued behind it. With enough of those parked at once, every worker is waiting on a future that cannot start. That is circular wait, not slow drain.

This is not theoretical. An unpatched 100-step delta sweep was run, hung, and was killed after five minutes. The one-line max_concurrency on the delta arm’s invoke config is a load-bearing workaround, not a precaution — and byte counts under it match unpatched runs at step counts where the unpatched run does finish, so the workaround does not perturb the measurement.

The LangGraph issue tracker was searched for this and nothing matching was found — first on 2026-08-30, again on 2026-10-08. That is a statement about the search, not about the bug’s reporting status.

Re-checked on 2026-10-08 against langgraph==1.2.14, three patch releases later. 1.2.13 shipped five DeltaChannel fixes, none of them this one: _checkpointer_put_after_previous is byte-for-byte unchanged (now pregel/_loop.py:1604-1606). The lab’s own delta graph, run at 60 steps with no max_concurrency, hung three times out of three; with the workaround it completed thirty times out of thirty. The lab’s full test suite passes on 1.2.14 unchanged. Its pin stays at 1.2.11, because that is the release its byte counts were measured on.

A non-associative reducer silently reconstructs different state. Replay folds writes in larger batches than they were produced in, so a reducer sensitive to batching returns something else after a resume. Measured on a 12-step run: 25 items live, 13 items after replay, no error raised anywhere. DeltaChannel requires a batching-invariant reducer and has no way to check that you supplied one — the failure is silent, and it is silently wrong state, not a crash.

Human approval resumes through interrupt() and Command(resume=...). The graph pauses, the run returns with __interrupt__ set, and the original compiled graph resumes it to completion.

Finished state survives into a completely fresh graph object. A second graph, compiled independently against the same saver and sharing nothing in-process with the first, reads the completed values back through get_state alone — no Command, no interrupt(), its node functions never executed.

Those are two pinned facts, and this page deliberately stops short of Module 7’s fuller claim. What the suite does not prove is that a still-pending interrupted checkpoint can be driven to completion by a fresh process with no Command. The crash-recovery arm reads finished state from a fresh object; it does not resume a pending one. That is a narrower result than “human-in-the-loop and crash recovery are one mechanism,” and the narrower result is the one the tests support.

Terminal window
cd labs/langgraph-checkpoint-cost
uv venv .venv && uv pip install --python .venv/bin/python -e '.[dev]'
./.venv/bin/python -m ruff check . && ./.venv/bin/python -m mypy src tests && ./.venv/bin/python -m pytest -q
./.venv/bin/python -m checkpoint_cost.cli

19 tests, ruff clean, mypy --strict clean. Unlike this repo’s other labs, mypy runs over tests as well as src: a chunk of this lab’s value is “trust these numbers,” so the code producing them is held to the same bar as the code under test. The last command prints the three byte tables above plus the resume sweep, whose raw timings are deliberately not published here.

  • The instrument is part of the result. This lab’s first table was arithmetically correct over an incomplete number, and it read as a finding. Ask of any cost measurement: what call sites does the meter wrap, and what does the system do that it never sees?
  • “The fix” is a state-model rewrite, and the beta risk is narrower than it sounds. Adopting DeltaChannel late means changing how state is modelled, under load, on a running system. The storage risk is the smaller half, but it is not zero: LangGraph says the API and the on-disk representation may change, and that threads already written are expected — not guaranteed — to remain readable. What it names outright as unstable is the surrounding contract, which bites hardest if you read delta state yourself, and the representation underneath get_state is itself only expected, not guaranteed, to stay readable. Know which of the two risks you are actually taking, and how firm the assurance under it is, before you quote either to a review.
  • A silent correctness failure beats a loud cost failure, every time. The quadratic write curve is visible on a dashboard. A non-associative reducer producing 13 items where 25 lived is visible to nobody.
  • Deferred cost is still cost, and it moves to a different owner. Write cost is paid by the service doing the work; resume cost is paid by whoever is waiting on get_state — often a user, often on the recovery path, which is the worst moment to discover it.
  • Module 7: LangGraph — state, reducers, and the checkpointing claims this lab measures.
  • LangGraph — the lookup whose DeltaChannel recommendation this lab qualifies.
  • Durable Agent Task Engine — checkpointed resumption and retry semantics as a system rather than as a channel.