AION
Research paperLarge Language Models · Efficiency & Inference · Multimodal Models2 sources · Oct 8, 2026

V-CoLA: Vision Token Compression with Linear Attention

To this end, we propose V-CoLA, an efficient training-free token compression framework specifically designed for linear attention.

Key points

  • Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence.
  • This motivates vision token compression as a key direction to alleviate the burden.
  • Our analysis reveals that both attention- and similarity-based approaches suffer notable performance degradation, underscoring the urgent need for compression methods tailored to this regime.
  • V-CoLA introduces a novel uniqueness-aware importance criterion for identifying critical vision tokens, coupled with an adaptive token merging strategy that performs compression.

Sources (2)

Extractive summary: sentences quoted from the sources.