ResearchResearch paperRetrieval, RAG & Search · Computer Vision1 source · Oct 7, 2026

Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval

Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers.

Key points

  • In this work, we augment the semantic tokens in the newer visual transformers, the global [CLS] token and the four register tokens, with a carefully selected collection of spatial tokens, aiming to capture the spatial region representation that characterizes the contents captured in each of the semantic tokens.
  • For each "cue" token ([CLS] and each register token), we find a "buddy" image patch token and extract an N x N patch region to produce a set of localized ROI tokens.
  • Furthermore, we incorporate these tokens into a multi-vector retrieval framework inspired by ColBERT, enabling fine-grained matching via a per-token alignment mechanism while avoiding the large storage cost of keeping all patch embeddings.
  • Through extensive experiments, we find that (1) register tokens encode useful fine-grained details that can complement the [CLS] token; (2) automatically pooled ROI tokens further improve fine-grained discrimination; and (3) multi-vector retrieval with a small set of tokens improves over a DINOv2-reg single-vector baseline while remaining tractable for large-scale search.

Sources (1)

  • [1]Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:24 PM
    Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers.
    In this work, we augment the semantic tokens in the newer visual transformers, the global [CLS] token and the four register tokens, with a carefully selected collection of spatial tokens, aiming to capture the spatial region representation that characterizes the contents captured in each of the semantic tokens.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Multimodal open d1 decision models for the edge
  2. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  3. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
  4. Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
  5. Oct 6, 2026Optimization Encoders: Rethinking Second-Order Meta-Learning for Neural Fields
  6. Oct 6, 2026huggingface/diffusers v0.41.0: Diffusers 0.41.0: QwenImage 2.1 pipeline and more

Related