Learning to Retrieve: Internalizing Memory Retrieval for Video World Models
We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system.
Key points
- Video world models aim to generate explorable, 3D-consistent scene videos conditioned on camera trajectories.
- Existing approaches often rely on external memory systems that explicitly retrieve previously observed content to mitigate scene drift during long-horizon generation.
- Based on this principle, we introduce Learning-to-Retrieve (L2R), which repurposes the model's persistent internal state as a memory for historical context.
- A camera-conditioned retrieval gate selectively accesses relevant historical information from this state, determining what to retrieve, while a retrieval trigger determines when to retrieve.
Sources (1)
- [1]Learning to Retrieve: Internalizing Memory Retrieval for Video World ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:01 AM
We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system.
Video world models aim to generate explorable, 3D-consistent scene videos conditioned on camera trajectories.
Extractive summary: sentences quoted from the sources.