V-CoLA: Vision Token Compression with Linear Attention
To this end, we propose V-CoLA, an efficient training-free token compression framework specifically designed for linear attention.
Key points
- Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence.
- This motivates vision token compression as a key direction to alleviate the burden.
- Our analysis reveals that both attention- and similarity-based approaches suffer notable performance degradation, underscoring the urgent need for compression methods tailored to this regime.
- V-CoLA introduces a novel uniqueness-aware importance criterion for identifying critical vision tokens, coupled with an adaptive token merging strategy that performs compression.
Sources (2)
- [1]V-CoLA: Vision Token Compression with Linear AttentionHugging Face Daily Papers · Oct 8, 12:00 AM
To this end, we propose V-CoLA, an efficient training-free token compression framework specifically designed for linear attention.
Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence.
- [2]V-CoLA: Vision Token Compression with Linear AttentionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:55 AM · same content
Extractive summary: sentences quoted from the sources.