RippleCP: Measuring Counterfactual Checkpoint Advantage in Coding Agents
Agent checkpoint systems decide what state is recovery-relevant, how to snapshot it, and whether rollback is admissible.
Key points
- None decides which of the safe boundaries they expose are worth materializing.
- We formulate this as counterfactual checkpoint advantage, the reduction in future recovery cost obtained by checkpointing a candidate rather than skipping it, and measure it by driving a CP branch and a SKIP branch to the same logical failure and recovering both under matched model, tool, verifier, and stopping conditions.
- On a frozen pilot of 12 SWE-bench Verified tasks and 106 real recovery branches, checkpointing saves 49.4 s per task, and that figure resolves into two regimes two orders of magnitude apart.
- Recovery is a re-derivation rather than a replay, so preserved work is a poor guide to saved work, and the classical elapsed-work rule misprices the second checkpoint by its full nominal cost.
Sources (1)
- [1]RippleCP: Measuring Counterfactual Checkpoint Advantage in Coding AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:40 PM
Agent checkpoint systems decide what state is recovery-relevant, how to snapshot it, and whether rollback is admissible.
None decides which of the safe boundaries they expose are worth materializing.
Extractive summary: sentences quoted from the sources.