AION
Research paperReinforcement Learning1 source · Oct 6, 2026

On KL-Regularized Policy Optimization

We propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler.

Key points

  • Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters.
  • Standard remedies either clip importance ratios, which biases the update, or, as in GRPO, sample a group of responses per prompt, which is costly when episodes are long.
  • For token-level policy mirror descent targets, we show that the resulting gradient can be computed from terminal returns without a critic, via sampler-centered scores or a single trajectory residual, even under stochastic tool outputs.
  • We further prove that independent Monte Carlo estimates of the KL term keep these gradients unbiased, derive the exact KL gap of cheaper top-$K$ and binary approximations, and show that SPPO, GPO, REBEL, and BPO arise as special cases of KLPO.

Sources (1)

  • [1]On KL-Regularized Policy Optimization
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:30 PM
    We propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler.
    Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters.

Extractive summary: sentences quoted from the sources.