AION
Research paperLarge Language Models · Safety & Alignment · Efficiency & Inference1 source · Oct 8, 2026

Right Screen, Wrong Transition: World Models as Verifiers for GUI Agents

We argue that a world model meant for verification should instead predict in the space in which observations are encoded, and present LGWM, a decoder-free, action-conditioned world model that predicts the representation of the next screen directly, trained without semantic annotation on 1.85M real GUI transitions.

Key points

  • A login screen that appears after a tap on Sign in is expected; the same screen after a tap on View order is an attack.
  • For GUI agents, safety is therefore a property of the transition rather than of the screen, and a monitor that inspects only screens can be defeated by reusing a legitimate one.
  • Existing GUI world models provide one, but they output it as text, code, or images, so checking it against the observed screen requires a second model to judge the two.
  • We evaluate on RSWT-BENCH, a diagnostic where each credential screen appears under both a legitimate and a hijacked transition, so detectors that see only the screen are at chance by construction.

Sources (1)

  • [1]Right Screen, Wrong Transition: World Models as Verifiers for GUI Agents
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:33 PM
    We argue that a world model meant for verification should instead predict in the space in which observations are encoded, and present LGWM, a decoder-free, action-conditioned world model that predicts the representation of the next screen directly, trained without semantic annotation on 1.85M real GUI transitions.
    A login screen that appears after a tap on Sign in is expected; the same screen after a tap on View order is an attack.

Extractive summary: sentences quoted from the sources.