Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution
To address this gap, we formulate red teaming for vision-grounded web agents as an end-to-end grounding-to-execution problem, and introduce WebMirage, a framework that crafts localized visual perturbations that cause agents to select attacker-controlled content and execute the corresponding browser action across varying webpage renderings.
ProofPaper ↗
Key points
- Modern web agents built on large vision-language models process webpages, select relevant UI elements, and translate model outputs into browser actions.
- Existing visual red-teaming approaches use adversarial visual content to manipulate this process.
- Consequently, model-level success does not establish control over browser execution and cannot reliably characterize end-to-end agent robustness.
- We evaluate WebMirage across four agent configurations and six VLM backbones on 2,250 tasks covering 13 public websites and a sandbox benchmark.
Sources (1)
- [1]Adversarial Images Hijack Web Agents from Visual Grounding to Browser ExecutionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:05 AM
To address this gap, we formulate red teaming for vision-grounded web agents as an end-to-end grounding-to-execution problem, and introduce WebMirage, a framework that crafts localized visual perturbations that cause agents to select attacker-controlled content and execute the corresponding browser action across varying webpage renderings.
Modern web agents built on large vision-language models process webpages, select relevant UI elements, and translate model outputs into browser actions.
Extractive summary: sentences quoted from the sources.