COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning
Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories.
Key points
- Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective.
- We show that this policy-side correction alone is insufficient: advantage estimates also inherit mismatch from behavior-policy continuations, which we term advantage staleness.
- We derive exact bias and variance decompositions for a general two-channel actor update, revealing nonseparable coupling between policy-weight and advantage-estimation errors: their interaction induces multiplicative bias terms, while squared policy weights amplify advantage uncertainty in gradient variance.
- We introduce Coupled Off-Policy Correction (COPC), an actor--critic method combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals for return and advantage estimation.
Sources (1)
- [1]COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement LearningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:43 AM
Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories.
Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective.
Extractive summary: sentences quoted from the sources.