ResearchResearch paperInterpretability · Large Language Models · Multimodal Models1 source · Oct 7, 2026

Detecting Adversarial Images through Response Profiles of Vision-Language Models

Adversarial perturbations can alter the predictions of frozen vision-language models (VLMs) while leaving their confidence and image--text similarity patterns seemingly plausible.

Key points

  • We investigate whether we can identify adversarial inputs based on the broader way an image interacts with a collection of general semantic prompts.
  • Our detector summarizes these responses using category-level statistics, relationships among prompts, deviations from clean reference distributions, and stability under weak image transformations, producing a compact response profile that is classified by a lightweight model while the VLM remains fixed.
  • We evaluate the approach on multiple public image datasets, several CLIP-style visual backbones, and a range of gradient-based, optimization-based, automated, and spatial attacks.
  • Overall, the results indicate that response patterns across semantic prompts provide a useful complementary signal for adversarial image detection in frozen VLMs.

Sources (1)

  • [1]Detecting Adversarial Images through Response Profiles of Vision-Language Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:13 PM
    Adversarial perturbations can alter the predictions of frozen vision-language models (VLMs) while leaving their confidence and image--text similarity patterns seemingly plausible.
    We investigate whether we can identify adversarial inputs based on the broader way an image interacts with a collection of general semantic prompts.

Extractive summary: sentences quoted from the sources.

Related