Learning in Dreams, Winning in Reality: A Continuous Dyna Loop for a Ten-Hero MOBA
We learn a structured, multi-agent world model of a complete ten-hero MOBA (206 units, every hero acting every tick, games of up to 6,000 ticks), train a policy only inside it with 1,400-tick free-running imagined episodes, and measure that policy in the real game against the opponent the game ships with.
Key points
- World models are usually judged from the inside: by prediction loss, by the return a policy earns in imagination, or by how convincing their frames look.
- Run as a continuous asynchronous Dyna loop, the policy wins 70.2% of real games as radiant (421 of 600; 95% CI 66.4-73.7) on seeds never used for any decision, up from 0% for dream training alone and 33.7% before the loop.
- Model exploitation is invisible from inside the dream: every unanchored run collapsed within a few updates while no in-dream metric tracked the collapse.
- We release the world model, the dream-PPO harness, a world-model debugger, the evaluation protocol, and every policy and log.
Sources (1)
- [1]Learning in Dreams, Winning in Reality: A Continuous Dyna Loop for a Ten-Hero MOBAarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 09:26 AM
We learn a structured, multi-agent world model of a complete ten-hero MOBA (206 units, every hero acting every tick, games of up to 6,000 ticks), train a policy only inside it with 1,400-tick free-running imagined episodes, and measure that policy in the real game against the opponent the game ships with.
World models are usually judged from the inside: by prediction loss, by the return a policy earns in imagination, or by how convincing their frames look.
Extractive summary: sentences quoted from the sources.