Detecting Adversarial Images through Response Profiles of Vision-Language Models
Adversarial perturbations can alter the predictions of frozen vision-language models (VLMs) while leaving their confidence and image--text similarity patterns seemingly plausible.
ProofPaper ↗
Key points
- We investigate whether we can identify adversarial inputs based on the broader way an image interacts with a collection of general semantic prompts.
- Our detector summarizes these responses using category-level statistics, relationships among prompts, deviations from clean reference distributions, and stability under weak image transformations, producing a compact response profile that is classified by a lightweight model while the VLM remains fixed.
- We evaluate the approach on multiple public image datasets, several CLIP-style visual backbones, and a range of gradient-based, optimization-based, automated, and spatial attacks.
- Overall, the results indicate that response patterns across semantic prompts provide a useful complementary signal for adversarial image detection in frozen VLMs.
Sources (1)
- [1]Detecting Adversarial Images through Response Profiles of Vision-Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:13 PM
Adversarial perturbations can alter the predictions of frozen vision-language models (VLMs) while leaving their confidence and image--text similarity patterns seemingly plausible.
We investigate whether we can identify adversarial inputs based on the broader way an image interacts with a collection of general semantic prompts.
Extractive summary: sentences quoted from the sources.