On KL-Regularized Policy Optimization
We propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler.
Key points
- Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters.
- Standard remedies either clip importance ratios, which biases the update, or, as in GRPO, sample a group of responses per prompt, which is costly when episodes are long.
- For token-level policy mirror descent targets, we show that the resulting gradient can be computed from terminal returns without a critic, via sampler-centered scores or a single trajectory residual, even under stochastic tool outputs.
- We further prove that independent Monte Carlo estimates of the KL term keep these gradients unbiased, derive the exact KL gap of cheaper top-$K$ and binary approximations, and show that SPPO, GPO, REBEL, and BPO arise as special cases of KLPO.
Sources (1)
- [1]On KL-Regularized Policy OptimizationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:30 PM
We propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler.
Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters.
Extractive summary: sentences quoted from the sources.