Foveated Compression: Selective High-Resolution Preservation for Token-Efficient VLMs
We introduce Foveated Compression, which encodes a full-resolution image once and represents it with a mixture of native- and compressed-resolution visual tokens.
ProofPaper ↗
Key points
- Visual tokens are a major source of inference cost in vision-language models, yet simple image downsampling remains a surprisingly strong compression baseline.
- This raises a complementary question: under a fixed token budget, where should visual fidelity be preserved?
- A behaviorally self-distilled Foveated Merger compresses local visual tokens while preserving compatibility with their native counterparts, and a lightweight Foveated Selector chooses one of nine spatial cells to retain at native resolution using exhaustive budget-matched intervention supervision.
- These results expose complementary bottlenecks in region selection and compressed-region fidelity.
Sources (1)
- [1]Foveated Compression: Selective High-Resolution Preservation for Token-Efficient VLMsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:28 AM
We introduce Foveated Compression, which encodes a full-resolution image once and represents it with a mixture of native- and compressed-resolution visual tokens.
Visual tokens are a major source of inference cost in vision-language models, yet simple image downsampling remains a surprisingly strong compression baseline.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
- Oct 6, 2026RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
- Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
- Sep 29, 2026Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
- Sep 28, 2026Notes on NVIDIA Nemotron
- Aug 10, 2026huggingface/transformers v5.15.0: Release: v5.15.0