ResearchResearch paperLarge Language Models · Multimodal Models · Efficiency & Inference1 source · Oct 8, 2026

Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning

To address this challenge, we propose ViMoD, a lightweight framework that maintains a compact visual context while preserving access to original fine-grained evidence as reasoning needs evolve.

Key points

  • Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs).
  • Existing one-shot pruning and aggregation methods compress visual tokens into a fixed context before decoding.
  • Temporal Routing for Adaptive Contextual Evidence (TRACE) integrates decoding history to anticipate upcoming evidence needs and select, retain, or replace active Fine-token groups.
  • Selected Fine tokens augment the persistent Coarse context in the frozen backbone, enabling stage-specific evidence access without continuously attending to all visual tokens.

Sources (1)

  • [1]Look Back, Think Ahead: Visual Memory on Demand for Efficient Multimodal Reasoning
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:42 PM
    To address this challenge, we propose ViMoD, a lightweight framework that maintains a compact visual context while preserving access to original fine-grained evidence as reasoning needs evolve.
    Processing long visual token sequences from high-resolution images makes multi-step reasoning computationally expensive for multimodal Large Language Models (MLLMs).

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026ConwayResearch/Underdog-Saluki-27B-1.0
  2. Oct 8, 2026SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
  3. Oct 8, 2026REMORY: Learning Residual Memory for Context Compaction
  4. Oct 8, 2026OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
  5. Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
  6. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

Related