AION
Research paperReinforcement Learning1 source · Oct 7, 2026

World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning

Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task.

Key points

  • Rather than deploying only the best-performing policy, we ask whether a set of frozen goal-conditioned policies can be used collectively as a portfolio, deciding at every state which policy should act.
  • The policies' own value functions cannot be compared directly: they may use different scales, and some policies have no value function.
  • WMPA assumes access to a bank of frozen goal-conditioned policies and requires neither policy retraining nor privileged task-specific knowledge.
  • Under the official OGBench evaluation protocol on 18 state-based datasets spanning maze navigation as well as cube, scene, and puzzle manipulation, WMPA improves the macro-average success rate from the 44% achieved by the best policy selected per dataset to 58%, with statistically significant gains on 12 datasets.

Sources (1)

  • [1]World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 09:41 PM
    Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task.
    Rather than deploying only the best-performing policy, we ask whether a set of frozen goal-conditioned policies can be used collectively as a portfolio, deciding at every state which policy should act.

Extractive summary: sentences quoted from the sources.