AION
Research paperLarge Language Models · Reinforcement Learning1 source · Oct 7, 2026

World Potential Model: Pretrained World Knowledge as Progress Potentials

We formalize this capability with a World Potential Model (WPM), a goal-conditioned evaluator of task-relative realized progress in agent contexts.

Key points

  • Long-horizon language agents often receive supervision only from terminal task outcomes, leaving little signal for distinguishing productive intermediate behavior from stagnation or even regression.
  • Rather than learning a separate value function or process reward model for every task, we ask whether pretrained models can recognize task progress from their existing world knowledge.
  • We further anchor these progress judgments to task-specific milestones to obtain scalar world potentials, whose temporal differences provide process-sensitive step-level credit for policy optimization.
  • Together, these results provide initial evidence that pretrained world knowledge can support reusable realized-progress evaluation and provide useful supervision for long-horizon agents.

Sources (1)

  • [1]World Potential Model: Pretrained World Knowledge as Progress Potentials
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:01 AM
    We formalize this capability with a World Potential Model (WPM), a goal-conditioned evaluator of task-relative realized progress in agent contexts.
    Long-horizon language agents often receive supervision only from terminal task outcomes, leaving little signal for distinguishing productive intermediate behavior from stagnation or even regression.

Extractive summary: sentences quoted from the sources.