AION
Research paperLarge Language Models · Interpretability · Multimodal Models1 source · Oct 8, 2026

Does Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive Encoders

Adversarial attacks on vision-language models optimize an image toward a text target, then cite the attacked model's similarity score as evidence of success.

Key points

  • We ask whether that score - victim-space target alignment (VTS) - predicts recovery of the target by an independent model.
  • Two preregistered studies then compare six contrastive encoders under a matched attack at three perturbation budgets.
  • Within robust encoders, per-sample alignment gain correlates with evidence gain ($ρ= 0.24-0.51$); within vanilla CLIP the correlation is consistent with zero.
  • Across encoders we find no monotone alignment-evidence relation.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related