World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning
Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task.
Key points
- Rather than deploying only the best-performing policy, we ask whether a set of frozen goal-conditioned policies can be used collectively as a portfolio, deciding at every state which policy should act.
- The policies' own value functions cannot be compared directly: they may use different scales, and some policies have no value function.
- WMPA assumes access to a bank of frozen goal-conditioned policies and requires neither policy retraining nor privileged task-specific knowledge.
- Under the official OGBench evaluation protocol on 18 state-based datasets spanning maze navigation as well as cube, scene, and puzzle manipulation, WMPA improves the macro-average success rate from the 44% achieved by the best policy selected per dataset to 58%, with statistically significant gains on 12 datasets.
Sources (1)
- [1]World-Model Policy Arbiter for Goal-Conditioned Reinforcement LearningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 09:41 PM
Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task.
Rather than deploying only the best-performing policy, we ask whether a set of frozen goal-conditioned policies can be used collectively as a portfolio, deciding at every state which policy should act.
Extractive summary: sentences quoted from the sources.