Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval
Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers.
ProofPaper ↗
Key points
- In this work, we augment the semantic tokens in the newer visual transformers, the global [CLS] token and the four register tokens, with a carefully selected collection of spatial tokens, aiming to capture the spatial region representation that characterizes the contents captured in each of the semantic tokens.
- For each "cue" token ([CLS] and each register token), we find a "buddy" image patch token and extract an N x N patch region to produce a set of localized ROI tokens.
- Furthermore, we incorporate these tokens into a multi-vector retrieval framework inspired by ColBERT, enabling fine-grained matching via a per-token alignment mechanism while avoiding the large storage cost of keeping all patch embeddings.
- Through extensive experiments, we find that (1) register tokens encode useful fine-grained details that can complement the [CLS] token; (2) automatically pooled ROI tokens further improve fine-grained discrimination; and (3) multi-vector retrieval with a small set of tokens improves over a DINOv2-reg single-vector baseline while remaining tractable for large-scale search.
Sources (1)
- [1]Region-Aware CLS Token Augmentation for Fine-Grained Image RetrievalarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:24 PM
Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers.
In this work, we augment the semantic tokens in the newer visual transformers, the global [CLS] token and the four register tokens, with a carefully selected collection of spatial tokens, aiming to capture the spatial region representation that characterizes the contents captured in each of the semantic tokens.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Multimodal open d1 decision models for the edge
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
- Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
- Oct 6, 2026Optimization Encoders: Rethinking Second-Order Meta-Learning for Neural Fields
- Oct 6, 2026huggingface/diffusers v0.41.0: Diffusers 0.41.0: QwenImage 2.1 pipeline and more