AION
Research paperLarge Language Models · Efficiency & Inference1 source · Oct 6, 2026

ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context.

Key points

  • Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window.
  • However, these predictive approaches introduce runtime overhead, invalidate prefix caches, and permanently discard content with no guarantee of recovery.
  • Both operators use chunked rendering, rewriting the cached prefix once every few steps rather than at every step.
  • Evaluations across five long-horizon benchmarks and two frontier LLMs demonstrate that ReFold reduces token consumption by up to 2.5x and halves the KV-cache memory per session without degrading task success rates.

Sources (1)

  • [1]ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:09 AM
    To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context.
    Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window.

Extractive summary: sentences quoted from the sources.