A Closed-Loop Non-Asymptotic Convergence Analysis of PPO with Learned Critics and Clipping
We develop a non-asymptotic analysis of PPO-Clip as a closed-loop actor--critic system.
Key points
- Despite its widespread use, Proximal Policy Optimization with clipping (PPO-Clip) remains difficult to tune, and the interactions among critic learning, clipping, and rollout reuse remain incompletely understood.
- It captures actor--critic coupling, nonsmooth probability-ratio clipping, finite-batch reuse, and predictable early stopping under explicit coverage and critic regularity assumptions, using raw GAE and Monte Carlo critic targets.
- Our synchronous and asynchronous guarantees jointly characterize policy stationarity and the tracking accuracy of the learned critic, with explicit dependence on algorithmic parameters.
- For finite layered MDPs with tabular critics, a uniform bound on the actual clipped-gradient class replaces complete-trajectory counting.
Sources (1)
- [1]A Closed-Loop Non-Asymptotic Convergence Analysis of PPO with Learned Critics and ClippingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:40 PM
We develop a non-asymptotic analysis of PPO-Clip as a closed-loop actor--critic system.
Despite its widespread use, Proximal Policy Optimization with clipping (PPO-Clip) remains difficult to tune, and the interactions among critic learning, clipping, and rollout reuse remain incompletely understood.
Extractive summary: sentences quoted from the sources.