AION
Research paperReinforcement Learning · Robotics & Embodied AI2 sources · Oct 7, 2026

System Switch: When Should a Fast Decision Model Stop and Think?

Dual-process agents pair a fast policy with a slow deliberative model.

Key points

  • In real-time settings the slow model usually runs continuously; in turn-based agents and robot planners it is invoked on events such as uncertainty or a detected failure.
  • We study a fast learned actor that takes every decision and hands control to a reasoning vision-language model only when a gate opens, while the game keeps running.
  • We use closed-loop Doom and the new open "System One" typed-decision models, served through a common llama.cpp interface.
  • We release code, prompts, data and logs.

Sources (2)

Extractive summary: sentences quoted from the sources.