AION
Research paperReinforcement Learning1 source · Oct 7, 2026

COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning

Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories.

Key points

  • Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective.
  • We show that this policy-side correction alone is insufficient: advantage estimates also inherit mismatch from behavior-policy continuations, which we term advantage staleness.
  • We derive exact bias and variance decompositions for a general two-channel actor update, revealing nonseparable coupling between policy-weight and advantage-estimation errors: their interaction induces multiplicative bias terms, while squared policy weights amplify advantage uncertainty in gradient variance.
  • We introduce Coupled Off-Policy Correction (COPC), an actor--critic method combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals for return and advantage estimation.

Sources (1)

  • [1]COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:43 AM
    Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories.
    Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective.

Extractive summary: sentences quoted from the sources.