ResearchResearch paperLarge Language Models · Interpretability1 source · Oct 7, 2026

Less from More: Reinforcing Sparse Video Reasoning from Dense References

We propose SAVER, a dense-to-sparse post-training framework that uses dense video views as training-time references for sparse-frame inference.

Key points

  • Video-language models commonly assume that more temporal observations lead to more reliable reasoning.
  • We question this assumption and argue that the key challenge is not merely processing more video frames efficiently, but learning to reason reliably under limited temporal evidence.
  • During reinforcement post-training, paired dense and sparse views are optimized with grounding rewards and a reliability-gated reference reward, encouraging sparse view predictions to preserve task-relevant temporal evidence.
  • These results show that temporal grounding can serve as an effective evidence-localization proxy for learning sparse video reasoning that transfers to broader video understanding tasks.

Sources (1)

  • [1]Less from More: Reinforcing Sparse Video Reasoning from Dense References
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:51 PM
    We propose SAVER, a dense-to-sparse post-training framework that uses dense video views as training-time references for sparse-frame inference.
    Video-language models commonly assume that more temporal observations lead to more reliable reasoning.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  2. Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
  3. Oct 6, 2026CARE: Certifying Acceleration for Vision-Language-Action Inference
  4. Oct 5, 2026perplexity-ai/pplx-decider-v1.1-27b
  5. Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
  6. Oct 4, 2026ausboss/Qwen-Image-2.1-Outfit-Swap-Consistency-LoRA

Related