Right Screen, Wrong Transition: World Models as Verifiers for GUI Agents
We argue that a world model meant for verification should instead predict in the space in which observations are encoded, and present LGWM, a decoder-free, action-conditioned world model that predicts the representation of the next screen directly, trained without semantic annotation on 1.85M real GUI transitions.
Key points
- A login screen that appears after a tap on Sign in is expected; the same screen after a tap on View order is an attack.
- For GUI agents, safety is therefore a property of the transition rather than of the screen, and a monitor that inspects only screens can be defeated by reusing a legitimate one.
- Existing GUI world models provide one, but they output it as text, code, or images, so checking it against the observed screen requires a second model to judge the two.
- We evaluate on RSWT-BENCH, a diagnostic where each credential screen appears under both a legitimate and a hijacked transition, so detectors that see only the screen are at chance by construction.
Sources (1)
- [1]Right Screen, Wrong Transition: World Models as Verifiers for GUI AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:33 PM
We argue that a world model meant for verification should instead predict in the space in which observations are encoded, and present LGWM, a decoder-free, action-conditioned world model that predicts the representation of the next screen directly, trained without semantic annotation on 1.85M real GUI transitions.
A login screen that appears after a tap on Sign in is expected; the same screen after a tap on View order is an attack.
Extractive summary: sentences quoted from the sources.