VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs
Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens.
ProofPaper ↗
Key points
- Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving.
- We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training.
- Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target.
- Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.
Sources (1)
- [1]VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:48 AM
Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens.
Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 2, 2026sgl-project/sglang v0.5.21
- Sep 18, 2026sgl-project/sglang v0.5.20
- Sep 5, 2026sgl-project/sglang v0.5.19
- Aug 22, 2026sgl-project/sglang v0.5.18
- Aug 8, 2026sgl-project/sglang v0.5.17
- Jun 13, 2026sgl-project/sglang v0.5.13