Skip to content

Architecture: Stateful Graph Checkpointing

Listen to this page17:05
Read the transcript

1. The run that gets slower as it goes

Host: So picture this: you build an agent, it flies through every test you throw at it in staging, and then you ship it. It runs fine for a while, and then someone notices a long-running session has gotten noticeably slower. Nothing in the code changed. What’s going on?

Guest: That’s actually a really common pattern, and it almost always traces back to how checkpointing works. Every graph runtime needs to persist state after each step so a crash or restart doesn’t lose the run, and the default way to do that is a full snapshot — serialize the whole state, write it, move on.

Host: That sounds reasonable though. Why would that get slower over time instead of just being a flat cost per step?

Guest: Because the state isn’t static — message history grows, retrieved documents pile up, tool results accumulate because later steps need them. If you rewrite the entire value every single step, and that value keeps growing every step, your total cost becomes quadratic in the length of the run. Each write still looks cheap in isolation, which is exactly why it’s invisible until someone runs the agent long enough for it to matter.

2. What a correct checkpointer actually has to guarantee

Host: Okay, so if the naive full-snapshot write is what’s killing you, what does a checkpointer actually have to guarantee to not have this problem? Let’s lay out the bar it has to clear.

Guest: First, it has to resume from the last committed step in a process that shares nothing in memory with the one that wrote it — no cheating with a live reference. Second, the per-step write cost has to be flat, not just smaller. A constant-factor speedup on a quadratic curve just buys you a few more steps before you hit the wall again, it doesn’t change the shape of the problem. And third, reading that state back has to have a bounded cost — reconstruction can’t mean walking all the way back to step zero every time.

Host: So resumability, flat writes, bounded reads — that’s the core three. What else is on the list?

Guest: Three more, and they’re the ones people skip. Reducer semantics have to survive reconstruction no matter how the read path batches writes back together — fold them wrong and you get a different state than what actually ran. You need an instrument that sees every path state gets written through, not just the obvious one, or you’ll miss writes and think you’re safe when you’re not. And because this store holds the full content of every run, you need an actual retention and classification policy, not just ‘it persists.’

3. The cost doesn’t disappear, it moves

Host: So if snapshotting is quadratic, the obvious fix is to stop snapshotting — just write the delta, the thing that changed, and move on. Why doesn’t that just solve it?

Guest: Because the cost doesn’t vanish, it relocates. Somebody still has to fold all those deltas back into a state before anything can resume, and now that work happens at read time instead of write time. You’ve handed the bill to whoever’s waiting on the resume, which is often the worst possible moment to hand it to them.

Host: And that reconstruction has to fold things back in some order. Is that order guaranteed to match what actually happened during the run?

Guest: No, and that’s the trap. The original run batched writes however the scheduler happened to batch them; replay just folds whatever ancestor writes it finds, in whatever grouping the reconstruction logic chooses. If a reducer’s output depends on how things were batched, resume gives you a different answer than what actually ran — and there’s no check for that, because the runtime has no idea your reducer cares about batching. Add to that the delta encodings are usually the newer, less-battle-tested code path on disk, the store still caps how big any one write can be, and persistence frequently shares the same worker pool as the actual graph execution — so the fix doesn’t remove the cost, it just moves it somewhere less visible and less tested.

4. Two strategies, one diagram

Host: So let’s actually draw the branch point, because I think people picture these two strategies as vaguely similar and they’re not. Walk me through what happens at a single step, full snapshot versus incremental.

Guest: Same superstep either way: a node returns a partial update, the reducer folds it into the channel, then the checkpointer persists. Incremental writes just that step’s delta, plus a full snapshot every N updates, so per-step cost stays flat and snapshot frequency is the dial — crank it high enough and you’ve effectively promised to replay the whole run on resume.

Host: Okay, flat per-step cost sounds like the obvious win then. What’s the part people skip past when they pick incremental?

Guest: One thing, and it’s invisible until something goes looking: reconstruction isn’t a read anymore, it’s deserializing the nearest snapshot and re-folding every write since through the reducer — which is exactly the re-fold correctness condition we were just talking about, and it’s the only piece on that whole diagram that fails silently.

5. Six ways this fails, and only one of them is loud

Host: Okay, you called the silent reconstruction failure the most dangerous item on the page, but let’s lay out the whole list, because I count at least six distinct ways this goes wrong. Start with the one everybody misses before they even ship: the quadratic curve itself.

Guest: Right, and it hides because nobody load-tests the axis that matters, which is step count. Short graphs just don’t reach the point where the curve separates from its linear cousin, because those tests run many short graphs instead of one long one.

Host: So the second failure is almost the opposite problem — you built the fix, but your instrument can’t even see it working.

Guest: Exactly, that one’s sneaky because it looks like good news. If your meter wraps the snapshot call, it reports near-zero bytes for the incremental strategy, because the state didn’t vanish, it moved to a writes table the meter never looks at. That’s not a cheap write, it’s a blind instrument producing a number that’s clean, plausible, and wrong — the lab’s own first measurement had exactly that defect before anyone caught it.

Host: And then there’s the ownership mismatch — write cost and resume cost don’t land on the same person, or at the same moment.

Guest: Right, and that mismatch means the incremental strategy is cheapest when things are going well and most expensive exactly when the system is already degraded. Pair that with checkpoints becoming a dumping ground for tool output and retrieved documents that never belonged there, and the per-step cost stops being set by the cursor size and starts being set by payload size, which makes the quadratic curve steeper by that whole factor — and under a shared executor, enough of those parked writes at once and you don’t get slow, you get a deadlock, which is at least loud.

6. The lab’s headline number: 22x, and a different shape

Host: So let’s put an actual number on this. The lab ran the same accumulating state through a plain reducer and through the incremental channel, same payloads, swept from 10 to 100 steps. What came out at the end?

Guest: At 100 steps the plain reducer is re-serializing about 1.34 megabytes through put, the incremental version is doing roughly 60 kilobytes once you add put and writes together. That’s over 22 times smaller, and it’s not a cherry-picked step — it’s the steady state after the cost has had room to build.

Host: But the headline number isn’t even the part you care most about, is it — you said the shape matters more than the constant.

Guest: Right, the constant just tells you who’s ahead today. The shape tells you who wins tomorrow — between 50 and 100 steps the plain arm’s put bytes roughly quadruple, 345K to 1.34 million, because it’s re-serializing an ever-growing list every time. The incremental arm just doubles, 30K to 60K, and a flat-cost control arm that replaces state instead of accumulating stays pinned at 538 bytes final-step regardless of step count — that’s what tells you the growth you’re seeing is accumulation, not LangGraph’s fixed per-checkpoint overhead.

7. The meter that measured the wrong thing

Host: So before you got those clean numbers, you got a wrong table first — and it looked fine. What broke?

Guest: The first version of the measuring saver only hooked put(). For most channels that’s the whole story, but DeltaChannel’s checkpoint() method returns MISSING — its committed state never lands in channel_values at all. So put() was faithfully totalling a sentinel, the checkpoint remainder, and some metadata, and calling that the per-step cost. It compiled, it ran, it produced plausible-looking numbers, and it was wrong, because the actual delta state only exists in the pending-writes table, written through put_writes(). We actually have a test that exists purely to fail against that blind instrument, checking that the delta channel’s state is captured by put_writes and not by put.

Host: Okay, so the fix is obvious then — just add put_writes to the meter and sum both columns.

Guest: That’s the trap, because summing is only correct for DeltaChannel. Every other channel’s put_writes call re-serializes the exact same node output that put() already counted for that step — add them and you’ve double-counted a quantity that was never duplicated in storage. DeltaChannel is the one case where the two numbers are disjoint, checkpoint and pending-writes genuinely hold different bytes, so there it’s the sum or nothing. A single total-bytes column can’t be right for both arms at once; you need to know which channel you’re measuring before you know whether to add or to pick one.

8. Where the saved cost actually goes: reading it back

Host: Okay, so DeltaChannel defers the write cost. But deferred isn’t free — it has to show up somewhere. Where does it land?

Guest: On read. Reconstructing state means replaying every ancestor write through the reducer, unless a full snapshot is sitting close behind to cut that chain short. snapshot_frequency is the dial that controls how long the chain gets, and the lab swept it at 5, 25, 100, and 10000 over a 100-step run — that last setting is big enough that no snapshot ever fires, so every resume there replays the entire run from scratch.

Host: So that’s the worst case, and presumably the number people want is how bad it actually gets.

Guest: And that’s exactly the number we don’t print to a precise decimal, on purpose — two earlier drafts tried and got burned by run-to-run variance. What twelve runs do support is two honest claims: every configuration resumes in tens to low hundreds of microseconds at 100 steps, and the never-snapshot setup is reliably the slowest of the four because it’s doing the deepest replay. Real, measurable, small at this scale — but that’s a statement about 100 steps, not a license to extrapolate to 10,000.

9. Beta storage, and a real deadlock in the pinned release

Host: Okay, so if DeltaChannel is the thing that actually earns its keep at scale, I assume the docs just say ‘use this, it’s fine’ — but you flagged it as beta. What’s the catch?

Guest: The catch is in the channel’s own docstring, which the reference page doesn’t quote. It says the API and the on-disk representation may change, and that existing threads are ‘expected’ to remain readable — not guaranteed. That’s a hedge about a format that’s still moving underneath you, not a compatibility promise.

Host: That’s a much smaller claim than ‘the fix for large-state runs.’ But you mentioned something sharper than beta instability — an actual deadlock?

Guest: Yeah, and it’s not the boring explanation. Node execution and checkpoint writes share one thread pool, but that alone doesn’t deadlock anything in a FIFO pool. The real bug is an order inversion — the delta write path checks its pending futures at execution time, not submission time, so a task can end up holding a worker while waiting on a put that was queued after it. I ran an unpatched 100-step delta sweep, it hung, and I killed it after five minutes. There’s a one-line concurrency cap that works around it, and I checked — the LangGraph issue tracker has nothing matching this, for whatever that’s worth.

10. 25 items live, 13 after replay, no error anywhere

Host: So that deadlock was loud — it hung, you killed it, you knew something was wrong. What’s the failure that doesn’t tell you anything’s wrong?

Guest: A non-associative reducer under replay. I ran a 12-step sequence, 25 items live at the end of the original run. Replayed the same checkpoints through DeltaChannel and got 13 items back — no exception, no warning, nothing in the logs. The reducer folds writes in whatever batch size replay happens to use, and if your reducer’s result depends on batch size, you silently get a different answer than the run that actually happened.

Host: And DeltaChannel has no way to catch that, because it doesn’t know what your reducer is supposed to guarantee. What does hold up, though — because you’ve also pinned two things that do work?

Guest: Resuming through interrupt and Command with resume works exactly as advertised — the graph pauses, comes back with interrupt set, the same compiled graph finishes it. And finished state really does survive into a totally fresh graph object, no shared process, just reading get state off the same saver. What I can’t say is the bigger claim people want — that a fresh process can drive a still-pending interrupt to completion with no Command. My tests only cover finished state landing in a fresh object, not a pending one being resumed by it. That’s a narrower result, and it’s the one I’ll actually defend.

11. The cheap fix before the hard one

Host: So if we zoom out from the narrow result you just defended — snapshots versus incremental — where does that actually leave someone deciding which one to run? It feels like you’ve spent this whole episode building the case for incremental, and now you’re going to tell me not to switch.

Guest: Pretty much, yes. That’s not a free upgrade, it’s a different bet.

Host: So before anyone touches the checkpointer itself, is there a cheaper move sitting underneath both of these options?

Guest: Almost always, yes. If the thing bloating your state is a large payload, store a reference to it instead of the payload itself, and every write at every step shrinks regardless of which strategy you’re on. Switching checkpointers is a state-model rewrite performed under load on a running system; trimming what goes into state is a local change you can make this afternoon, so try that first and only reach for the strategy change when accumulation is actually the thing your graph is for.

12. The checkpoint store is a liability, and the checklist to run before shipping

Host: So we’ve spent this whole episode on the cost curve, but I want to end somewhere I think people skip past: the checkpoint store itself. What is it, really, from a security standpoint?

Guest: It’s the run. Every message, every retrieved document, every tool argument, stored in full, inheriting the highest data classification of anything that passed through the graph. Teams provision it like a cache and treat deletion like it’s one row, but under an incremental strategy a run is a snapshot plus a chain of writes, so removing the checkpoint and leaving the write log hasn’t deleted the data, it’s deleted the index to it. And a checkpoint is a resumption capability — anyone who can read one can hand it to the runtime and continue that run as that run, with everything it’s accumulated, so read access is execution access, and it’s rarely scoped like that.

Host: And tenant identity needs to be structural, not incidental, because reconstruction is walking ancestors by thread — a walk that crosses a tenant boundary is a much worse bug than a slow one. So if someone’s shipping this tomorrow, what’s actually on the checklist?

Guest: Know your step-count distribution instead of assuming it, put per-channel bytes and the first-to-last ratio on a dashboard, hold large payloads by reference before you touch the persistence strategy at all, and make every reducer you pair with incremental batching-invariant with a real test. Reconstruct a finished run against its live state on a schedule, set the snapshot frequency setting on purpose rather than by default, and measure reconstruction latency on the recovery path, not the happy one. None of that is exotic — it’s just measuring your own state model instead of borrowing someone else’s numbers, which is the whole bet this episode has been about.

Not covered

The planner wanted these and found nothing in the source to support them:

  • A direct side-by-side cost comparison with durable-agent-execution’s checkpoint-size guidance as a unified best-practice
  • Semantic caching’s false-hit framing as a general analogy for checkpointing risk beyond the single cost-category comparison already stated in the source

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

A multi-step agent has to survive the process running it. The mechanism every graph runtime reaches for is the same: after each step, write the state somewhere durable, so a later process can read it back and carry on. Durable Agent Execution covers the delivery layer around that — leases, fencing, retries. This page covers the write itself.

The default implementation is a full snapshot per step: serialize the current value of each channel that changed, store it, move on. That is correct, simple, and fine — right up to the point where the state accumulates.

Accumulation is the normal case, not an edge case. Message history grows. Retrieved documents collect. Tool results pile up because a later step needs an earlier one. And when the value grows every step and the whole value is rewritten every step, the cost of a run is quadratic in its length while every individual write still looks cheap.

That shape is the problem. It does not show up in tests, which run short graphs. It does not show up as an error. It shows up as a run that gets slower as it progresses, on a system that was fast in staging, and the profile points at storage rather than at the state model that is actually responsible.

  • Resume from the last committed step, in a process that shares nothing in-memory with the one that wrote it.
  • Per-step write cost that does not grow with step count. Not “small” — flat. A constant-factor improvement to a quadratic curve buys a fixed number of extra steps and nothing else.
  • Bounded reconstruction. Reading state back must have a ceiling that does not walk to step zero.
  • Reducer semantics preserved across reconstruction, whatever batching the read path happens to fold writes in.
  • An instrument that sees every path state is written through, not the most obvious one.
  • A retention and classification policy for the store, because it holds the full content of every run.
  • A snapshot write is a rewrite, not a diff. The cost of step n is proportional to the state at step n, so the cost of the run is the sum of that series. Linear growth per step is quadratic over the run.
  • The cost cannot be removed, only moved. Writing deltas instead of snapshots does not make persistence free; it defers reconstruction to whoever reads the state back. Deferred cost is still cost, and it changes owner — from the service doing the work to whoever is waiting on the resume.
  • Replay does not preserve the batching that produced the writes. A reconstruction folds ancestor writes in whatever batches it finds them in. Any reducer whose result depends on batching returns something different after a resume, and no runtime can check that yours does not.
  • Incremental encodings are younger than snapshot ones. The delta path in a given runtime is usually the newer, less-exercised code, often with an explicitly unstable on-disk representation — a real consideration when the data outlives the release that wrote it.
  • The store bounds the checkpoint. Row-size limits, request-size limits, and serialization budgets all put a ceiling on what one step may write, and a growing state reaches it eventually.
  • Persistence and execution often share an executor. Where checkpoint writes are submitted into the same bounded pool that runs the nodes, persistence work and useful work contend for the same workers — and can do worse than contend, as the failure modes below show.

The diagram is one superstep and its consequences. A node returns a partial update; the reducer folds it into the channel; the checkpointer persists the result. The branch is the whole design.

Full snapshot re-serializes the entire channel value. Every step pays for all the state that exists, so the run pays the sum of the series.

Incremental writes only the step’s own update, plus a full snapshot every N updates. Per-step cost is flat. What changes is the read: reconstruction deserializes the nearest snapshot and then replays every write since, through the reducer. snapshot_frequency is the dial between write volume and replay depth, and setting it high enough that no snapshot is ever written means every resume replays the entire run.

Two things the diagram makes explicit, because both are routinely missed:

Where the state actually lives changes. Under the incremental strategy the checkpoint may hold only a sentinel, with the real per-step state in a separate writes table. Anything that measures, audits, or exports “the checkpoint” is now looking at the wrong place.

Replay is a re-fold, not a read. That is where the correctness condition enters, and it is the only one on this page that fails silently.

Quadratic write cost, discovered in production

Short graphs hide it completely — at ten steps the difference is a few kilobytes and nobody looks. The curve only separates from its linear cousin at the step counts real runs reach. In the backing lab, doubling the run from 50 to 100 steps roughly quadrupled the accumulating channel’s total write volume while merely doubling the incremental arm’s. The symptom is “long runs get slower”, which gets attributed to the model, the network, or the database before anyone suspects the state model.

A reducer that is not batching-invariant

Replay folds writes in larger batches than execution produced them, so a reducer sensitive to batching reconstructs a different value. The backing lab measured this on a twelve-step run: 25 items live, 13 items after replay, and no error raised anywhere. No exception, no log line, no metric — the run simply continues from state that is not the state it had. This is the most dangerous item on the page, because every other failure here is loud.

An instrument that only watches one call path

When state moves to a writes table, a meter wrapped around the snapshot call reports near-zero bytes for the incremental strategy. That is not a cheap write, it is a blind instrument, and the number it produces is clean, plausible, and wrong. The lab’s first measurement had exactly this defect and produced a table that read as a finding; the test that now fails against a blind instrument exists because of it.

Resume cost paid on the recovery path

Write cost is paid by the service doing the work, during work it was already doing. Reconstruction cost is paid by whoever calls for the state — often a user, often immediately after a crash. The incremental strategy is therefore cheapest exactly when things are going well and most expensive at the moment the system is already degraded.

The checkpoint used as a result store

A checkpoint is a resumption cursor. Once intermediate results, retrieved documents, or raw tool output are parked in it because state was the convenient place to put them, per-step cost is set by the size of the payload rather than by the size of the cursor — and the quadratic curve is steeper by that whole factor.

Persistence contending with execution in one bounded pool

Checkpoint writes submitted into the same executor as node work do not merely queue; they can invert. In the release the backing lab pins, a checkpoint task drains and waits on the set of pending delta writes as found at execution time — which can include writes submitted after that task was. Enough of those parked at once and every worker is blocked on a future that cannot start. An unpatched 100-step sweep hung and was killed after five minutes.

Never snapshotting

Setting the snapshot interval beyond the length of a run removes snapshot writes entirely and makes replay depth equal to run length. It is the cheapest configuration to write and the slowest to read, and it is the one a benchmark that only measures writes will recommend.

  • Step count is the axis nobody load-tests. Throughput tests run many short graphs; the failure here needs one long one. A load profile that never exceeds ten steps cannot distinguish a linear cost curve from a quadratic one.
  • State size per step multiplies the whole curve. Halving the payload halves every write at every step. Keeping large values out of state entirely — a reference to object storage rather than the object — is the change with the best ratio of effort to effect, and it needs no new persistence strategy.
  • snapshot_frequency trades write volume against replay depth, and the two move in opposite directions. Frequent snapshots approach the snapshot strategy’s write cost; rare ones approach an unbounded read.
  • Reconstruction cost is real and, at these scales, small. Across the lab’s twelve runs every configuration reconstructed a 100-step run in tens to low hundreds of microseconds. That is the honest size of the deferred cost at 100 steps on an unloaded machine — enough to rule it out as a reason not to adopt incremental writes, and not enough to extrapolate to 10,000 steps.
  • The store’s per-row ceiling arrives before the cost curve becomes unbearable. A growing snapshot eventually exceeds a limit and the run fails outright, which is at least a loud failure.
  • Concurrency is bounded by the executor, not by the storage. Adding storage throughput does not help when the constraint is workers in a shared pool.
  • The checkpoint store holds the run. Every message, every retrieved document, every tool argument, in full. It inherits the highest data classification of anything that passes through the graph, and it is routinely provisioned as though it were a cache.
  • Deletion is not one row under an incremental strategy. A run’s state is a snapshot plus a chain of writes, potentially across tables. A deletion or export request that removes the checkpoint and leaves the write log has removed the index, not the data.
  • A checkpoint is a resumption capability. Anyone who can read one and hand it to the runtime can continue that run — as that run, with its accumulated context. Read access to the store is execution access to the workload, and it is rarely scoped as though it were.
  • Accumulated state accumulates personal data. The value that makes the write curve quadratic is usually conversation history, so retention on this store is a privacy control and not only a cost control.
  • Tenant identity belongs in the key, structurally. Reconstruction walks ancestors by thread; a walk that can cross a tenant boundary is a much worse bug than a slow one.

Full snapshot vs. incremental writes

Snapshots are simple, have no replay semantics to get wrong, and reconstruct in one read. They are quadratic over an accumulating run. Incremental writes are flat per step — the backing lab measured roughly 1.34 MB against roughly 60 KB over 100 steps, over 22× — and buy that with a reducer condition the runtime cannot verify, a deeper read path, and usually a younger storage format. For short runs or non-accumulating state, snapshots are not a compromise; they are the right answer.

Changing the persistence strategy vs. changing the state model

The cheapest fix is usually not a different checkpointer. State that does not accumulate has a flat curve under the default strategy, and holding a reference instead of a payload shrinks every write at every step. Switching strategies is a state-model rewrite performed under load on a running system; trimming what goes into state is a local change. Try the local change first, and reach for the strategy when the accumulation is the thing the graph is actually for.

snapshot_frequency: write volume vs. replay depth

Every snapshot written is a bounded replay later. Tune it by the cost you can least afford to pay at the wrong moment — and remember that the resume is usually the recovery path, so a design that optimises writes at the expense of reads has made the bad day worse to save money on the good one.

Adopting a beta incremental channel

The measured write-path effect is real and large. Against it: an on-disk representation that the implementation may change, a surrounding contract that is explicitly not yet stable, and a write path with less production exposure than the default. The storage risk is usually the smaller half — runtimes tend to commit to keeping already-written threads readable — but “expected to remain readable” is a stated expectation, not a guarantee, and the two are worth telling apart before either is quoted in a review.

Instrument fidelity vs. simplicity

A meter on one call site is easy to write and easy to trust. A meter that covers every path state can be written through requires knowing the runtime’s internals and produces columns that cannot honestly be summed. The second is more work and is the only one that answers the question.

  • The write curve is the headline, and its shape matters more than its constant. Over 100 steps the backing lab measured roughly 1.34 MB for a plain accumulating channel against roughly 60 KB for the same accumulation written incrementally. Extending the run widens that gap superlinearly.
  • Most of that cost is not the storage bill. Bytes are cheap. What is expensive is the serialization CPU on the critical path of every step, and the latency it adds to a run that a user may be waiting on.
  • Reconstruction is the cost you have chosen to pay later, on a path where latency is least welcome. At the scales measured it is sub-millisecond; the point is that it is not zero and that it grows with replay depth.
  • The engineering cost of switching is the real number in most cases. Re-modelling state on a running system, with a reducer condition to prove and a beta format to accept, is a larger expenditure than the bytes it saves — until the run lengths make it the only option.
  • The cost of getting it wrong silently is unbounded and does not appear on any invoice. A reconstruction that returns 13 items where 25 lived costs whatever a wrong answer costs, which is the same category of expense as a false cache hit in Semantic Response Caching.

The numbers here are a shape, not a benchmark for your system

The backing lab measures one runtime at a pinned version, with a fixed payload, no model calls, and disk deliberately excluded — which is what makes its byte counts reproduce exactly. Two things travel from it: the curve is quadratic against linear, and an instrument that watches one call path will miss where incremental state lives. The absolute byte counts do not travel, and neither does the lab’s refusal to rank the incremental strategy against its flat-cost control: under three accountings the incremental arm came out lower and under a fourth it came out 7% higher, and a gap that changes sign with the accounting rule is a property of the rule. Measure your own state model, and report what your meter does not see.

  • Bytes serialized per checkpoint, per channel, per step. One number for the whole write hides which channel is growing, and the growing one is the finding.
  • Bytes at the final step versus the first. The ratio is the accumulation, stated directly. A flat ratio means the state model is not the problem, whatever the total looks like.
  • Every write path, counted separately and never pre-summed. Where the same step’s output passes through two calls, adding them double-counts; where it passes through only one, omitting the other undercounts. Separate columns are the only honest presentation.
  • Step-count distribution of real runs. This is what says whether the quadratic term matters yet, and it is usually not collected at all.
  • Reconstruction latency, at percentiles, tagged by whether it was a recovery. The mean is dominated by warm resumes and says nothing about the case that hurts.
  • Replay depth — steps since the last snapshot — as a distribution. It is the input to reconstruction cost, and it is observable before anyone is waiting on it.
  • Checkpoint write latency against node latency in the same pool. When persistence starts competing with execution, this is where it appears first.
  • A reconstruction equality check, sampled. Reconstruct a finished run and compare against the state it ended with. It is the only signal that catches a batching-sensitive reducer, and nothing else on this list will ever fire for it.

Before long runs reach real traffic

  • The step-count distribution of real runs is known, not assumed.
  • Per-step serialized bytes are recorded per channel, and the first-to-last ratio is on a dashboard.
  • The instrument’s coverage is written down: which call sites it wraps, and what the runtime does that it never sees.
  • Large payloads are held by reference rather than carried in state, before any persistence strategy is changed.
  • Every reducer used with an incremental strategy is batching-invariant, with a test that folds the same writes in different batch sizes and asserts the same result.
  • A finished run is reconstructed and compared against its live final state, on a schedule.
  • snapshot_frequency is set deliberately, and the never-snapshot configuration is ruled out explicitly rather than by default.
  • Reconstruction latency is measured on the recovery path, not only on the happy one.
  • Checkpoint writes and node execution have been observed under concurrency, with the executor’s bound known.
  • The store’s retention, classification, and deletion semantics cover the write log as well as the checkpoint.
  • Read access to the checkpoint store is scoped as execution access, because that is what it is.

Hands-on Lab

A running implementation against a pinned langgraph==1.2.11: three graphs differing only in how state accumulates, a checkpointer wrapper that counts bytes on every path state is written through, a resume sweep across snapshot frequencies, and tests pinning the batching-invariance failure and a deadlock in the pinned release’s incremental write path. It also records what it retracted — a first table produced by an instrument that watched one call site. Read the lab documentation →

labs/langgraph-checkpoint-costproduction-shaped

A graph run gets slower with every step. Where do you look first?

At whether state accumulates and whether the checkpointer rewrites all of it each step. If both are true the run is quadratic in its own length by construction, and no amount of storage tuning changes the exponent. The confirming measurement is serialized bytes per step: flat means look elsewhere, growing means the state model is the cause.

Incremental checkpointing cut write volume 22×. What did you buy it with?

Read cost and a correctness condition. Reconstruction now deserializes the nearest snapshot and replays every write since, so the cost moved to whoever calls for the state — usually on the recovery path. And replay folds those writes in different batches than execution produced them, so the reducer has to be batching-invariant, which the runtime cannot check for you. The write saving is real; describing it as free is the error.

After a resume, a run has 13 items where 25 were live. No errors. What happened?

The reducer is not batching-invariant. Replay folded the ancestor writes in larger batches than the original execution did, and a fold whose result depends on batching returns something else. It is silent by nature: nothing was lost at write time and nothing threw at read time, so the only way to catch it is to reconstruct a finished run and compare against the state it ended with.

Your incremental write path measures near zero bytes per step. Do you believe it?

Not without knowing what the meter wraps. If the strategy stores a sentinel in the checkpoint and the real state in a separate writes table, a meter on the checkpoint call reports almost nothing and the bytes are simply elsewhere. The two paths also cannot be added blindly — for a channel whose whole value already round-trips through the checkpoint, summing them double-counts the same step’s output.

What is the cheapest fix for an expensive checkpoint?

Stop putting expensive things in it. A checkpoint is a resumption cursor; large payloads belong in object storage with a reference in state. That shrinks every write at every step, needs no new persistence strategy, and is a local change rather than a state-model rewrite on a running system. Switching strategies is the answer when the accumulation is the workload.

How would you set the snapshot interval?

By deciding which cost you can least afford at the wrong moment. Each snapshot bounds a later replay, so frequent snapshots move cost back onto the write path and rare ones deepen every resume. The configuration worth naming explicitly is the one where no snapshot is ever written inside a typical run — it is the cheapest to write, the slowest to read, and it is what a benchmark that measures only writes will recommend.