Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning
To address this challenge, we propose ViMoD, a lightweight framework that maintains a compact visual context while preserving access to original fine-grained evidence as reasoning needs evolve.
ProofPaper ↗
Key points
- Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs).
- Existing one-shot pruning and aggregation methods compress visual tokens into a fixed context before decoding.
- Temporal Routing for Adaptive Contextual Evidence (TRACE) integrates decoding history to anticipate upcoming evidence needs and select, retain, or replace active Fine-token groups.
- Selected Fine tokens augment the persistent Coarse context in the frozen backbone, enabling stage-specific evidence access without continuously attending to all visual tokens.
Sources (1)
- [1]Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal ReasoningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:42 PM
To address this challenge, we propose ViMoD, a lightweight framework that maintains a compact visual context while preserving access to original fine-grained evidence as reasoning needs evolve.
Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs).
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026ConwayResearch/Underdog-Saluki-27B-1.0
- Oct 8, 2026SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
- Oct 8, 2026REMORY: Learning Residual Memory for Context Compaction
- Oct 8, 2026OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
- Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning