AION
Research paperReinforcement Learning · Robotics & Embodied AI1 source · Oct 8, 2026

CausalDreamer: Learning Predictive World Models with Latent Disentanglement

We propose CausalDreamer, which keeps the tokenizer frozen and re-encodes its latent into a factored representation of four groups along two axes: controllability, where only the two controllable groups receive the action, and reward relevance, learned by predicting the reward from the two reward-relevant groups.

Key points

  • World models for control must capture which aspects of the environment respond to the agent's actions and which are relevant to reward.
  • Generative world models such as Dreamer 4 consist of a video tokenizer, which encodes each frame into a latent, and a dynamics model, which is pretrained to predict future latents from past latents and actions.
  • The pretrained dynamics model is then fine-tuned to predict the factored representation.
  • We evaluate CausalDreamer and the pretrained world model it starts from with model-predictive planning on 20 MMBench2 tasks: 10 clean tasks seen during training and 10 unseen tasks, of which 6 are manipulated variants of clean tasks with a changed background, object, or maze layout, and 4 are new environments.

Sources (1)

  • [1]CausalDreamer: Learning Predictive World Models with Latent Disentanglement
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:12 PM
    We propose CausalDreamer, which keeps the tokenizer frozen and re-encodes its latent into a factored representation of four groups along two axes: controllability, where only the two controllable groups receive the action, and reward relevance, learned by predicting the reward from the two reward-relevant groups.
    World models for control must capture which aspects of the environment respond to the agent's actions and which are relevant to reward.

Extractive summary: sentences quoted from the sources.