Memory Forcing: Attendable Mid-Horizon History for Streaming Video Generation
Autoregressive video diffusion enables causal video streaming without a bidirectional pass over the full clip, but existing few-step systems usually retain only the opening and most recent frames in a fixed-size KV cache.
Key points
- Once an event leaves this window, later frames can no longer attend to it, a failure we term mid-horizon forgetting.
- We present Memory Forcing, a few-step streaming method that preserves this missing history without increasing the cache size.
- Because absolute temporal indices drift outside the training range, Bank-aware RoPE reassigns indices at attention time so each bank remains distinguishable.
- At 1.3B, Memory Forcing leads on longer clips, shows the smallest drop from 5s to 60s among methods reporting all four lengths, and preserves subjects and scenes through leave-and-return.
Sources (1)
- [1]Memory Forcing: Attendable Mid-Horizon History for Streaming Video GenerationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 11:42 AM
Autoregressive video diffusion enables causal video streaming without a bidirectional pass over the full clip, but existing few-step systems usually retain only the opening and most recent frames in a fixed-size KV cache.
Once an event leaves this window, later frames can no longer attend to it, a failure we term mid-horizon forgetting.
Extractive summary: sentences quoted from the sources.