AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model.
Key points
- Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal.
- Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them.
- Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it.
- Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser.
Sources (1)
- [1]AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World ModelarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:56 PM
We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model.
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal.
Extractive summary: sentences quoted from the sources.