ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents
To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context.
Key points
- Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window.
- However, these predictive approaches introduce runtime overhead, invalidate prefix caches, and permanently discard content with no guarantee of recovery.
- Both operators use chunked rendering, rewriting the cached prefix once every few steps rather than at every step.
- Evaluations across five long-horizon benchmarks and two frontier LLMs demonstrate that ReFold reduces token consumption by up to 2.5x and halves the KV-cache memory per session without degrading task success rates.
Sources (1)
- [1]ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:09 AM
To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context.
Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window.
Extractive summary: sentences quoted from the sources.