ResearchResearch paperMultimodal Models · Large Language Models1 source · Oct 6, 2026

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens.

Key points

  • Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving.
  • We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training.
  • Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target.
  • Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.

Sources (1)

  • [1]VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:48 AM
    Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens.
    Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 2, 2026sgl-project/sglang v0.5.21
  2. Sep 18, 2026sgl-project/sglang v0.5.20
  3. Sep 5, 2026sgl-project/sglang v0.5.19
  4. Aug 22, 2026sgl-project/sglang v0.5.18
  5. Aug 8, 2026sgl-project/sglang v0.5.17
  6. Jun 13, 2026sgl-project/sglang v0.5.13

Related