Unsupervised Long-Tailed Adaptation of Vision-Language Models
Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data.
Key points
- To tackle this, we formalize a new scenario termed Unsupervised Long-Tailed Adaptation (ULTA).
- To address these issues, we propose a novel model called Margin-Aware Refinement with Structural alignment (MARS).
- Specifically, we mitigate head-class boundary erosion via Boundary-Preserving Alignment, which takes the zero-shot VLM as a fixed visual reference to suppress probability increases that lack visual support in the training targets.
- Building upon this, we introduce Margin-aware Self-Refinement, which employs a dynamic adjustment strategy to refine tail and confusable classes while preventing prediction bias.
Sources (1)
- [1]Unsupervised Long-Tailed Adaptation of Vision-Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:49 AM
Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data.
To tackle this, we formalize a new scenario termed Unsupervised Long-Tailed Adaptation (ULTA).
Extractive summary: sentences quoted from the sources.