AION
Research paperRobotics & Embodied AI1 source · Oct 8, 2026

WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models

World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT).

Key points

  • In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key-value (KV) pairs that the action expert queries.
  • We present WAM-Cache, a training-free framework that retains layerwise key-value representations across chunks and recomputes only a sparse refresh set of tokens.
  • Crucially, we find that the intuitive heuristic of refreshing visually drifted tokens plateaus far below the dense baseline, even with an oracle predicting ground-truth KV drift.
  • On Fast-WAM, WAM-Cache cuts video DiT prefill FLOPs by 32-42% across RoboTwin 2.0, LIBERO, and real-world experiments, while staying within 0.7-1.8 percentage points of the dense policy in simulation and 2.5 points on a real robot.

Sources (1)

  • [1]WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 07:33 AM
    World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT).
    In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key-value (KV) pairs that the action expert queries.

Extractive summary: sentences quoted from the sources.