AION
Research paperMultimodal Models · Large Language Models · Interpretability1 source · Oct 8, 2026

Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder

Building on this finding, we propose ComCLIP, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2.

Key points

  • CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA.
  • Post-training offers a lightweight route to refine CLIP, but recent work argues that the standard contrastive loss is unsuitable for post-training due to catastrophic forgetting under small batches, motivating designs that abandon the contrastive objective in favor of distillation.
  • We revisit this premise and find that, for the InfoNCE objective, the reported forgetting is driven primarily not by insufficient negatives but by an inappropriate magnitude of the contrastive temperature $τ$: with $τ$ set sufficiently small, contrastive post-training improves rather than degrades the pretrained CLIP, which we explain through the temperature dependence of the InfoNCE gradient.
  • Over multiple seeds, ComCLIP matches the self-distillation baseline CLIP-Refine on zero-shot classification while significantly improving the transferability of visual features, measured by linear probing ($48.99$ vs.\ $42.28$ on ViT-B/16), and on ViT-L/14 it also improves MMVP over CLIP-Refine ($24.20$ vs.\ $19.01$); CLIP-Refine remains stronger on image-text retrieval.

Sources (1)

  • [1]Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 07:06 AM
    Building on this finding, we propose ComCLIP, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv
    CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA.

Extractive summary: sentences quoted from the sources.