AION
Research paperReinforcement Learning · Training & Scaling1 source · Oct 7, 2026

A Closed-Loop Non-Asymptotic Convergence Analysis of PPO with Learned Critics and Clipping

We develop a non-asymptotic analysis of PPO-Clip as a closed-loop actor--critic system.

Key points

  • Despite its widespread use, Proximal Policy Optimization with clipping (PPO-Clip) remains difficult to tune, and the interactions among critic learning, clipping, and rollout reuse remain incompletely understood.
  • It captures actor--critic coupling, nonsmooth probability-ratio clipping, finite-batch reuse, and predictable early stopping under explicit coverage and critic regularity assumptions, using raw GAE and Monte Carlo critic targets.
  • Our synchronous and asynchronous guarantees jointly characterize policy stationarity and the tracking accuracy of the learned critic, with explicit dependence on algorithmic parameters.
  • For finite layered MDPs with tabular critics, a uniform bound on the actual clipped-gradient class replaces complete-trajectory counting.

Sources (1)

  • [1]A Closed-Loop Non-Asymptotic Convergence Analysis of PPO with Learned Critics and Clipping
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:40 PM
    We develop a non-asymptotic analysis of PPO-Clip as a closed-loop actor--critic system.
    Despite its widespread use, Proximal Policy Optimization with clipping (PPO-Clip) remains difficult to tune, and the interactions among critic learning, clipping, and rollout reuse remain incompletely understood.

Extractive summary: sentences quoted from the sources.