Unlocking Fine-Grained Perception in CLIP via Structurally-Aware Latent Masked Modeling
Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities.
ProofPaper ↗
Key points
- Existing research has attempted to enhance CLIP's visual representations by incorporating geometric priors from vision-centric models.
- To address these limitations, we propose SALM, an unsupervised embedding alignment framework based on structurally-aware latent mask modeling.
- First, we introduce a dual-matrix alignment strategy that explicitly calibrates intra-sample spatial correlations and activation intensities, thereby effectively injecting local geometric priors.
- Based on this, we further design a latent mask modeling mechanism to guide CLIP to restore the missing semantic details of the target model, thereby implicitly aggregating fine-grained structures into the global semantic space.
Sources (1)
- [1]Unlocking Fine-Grained Perception in CLIP via Structurally-Aware Latent Masked ModelingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:25 AM
Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities.
Existing research has attempted to enhance CLIP's visual representations by incorporating geometric priors from vision-centric models.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
- Oct 6, 2026RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
- Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
- Sep 29, 2026Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
- Sep 28, 2026Notes on NVIDIA Nemotron
- Aug 10, 2026huggingface/transformers v5.15.0: Release: v5.15.0