AION
Research paperReinforcement Learning1 source · Oct 6, 2026

Convex-Concave Reinforcement Learning

Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today.

Key points

  • Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even under a direct policy parameterization, and the field has largely responded by avoiding it: optimizing convex surrogate approximations of the return under trust-region constraints (NPG, TRPO, PPO, AWR).
  • In log-density-ratio coordinates $y := \log[π/πn]$, the exact per-iteration objective, computable via per-decision importance sampling (PDIS), is a difference-of-convex-constrained difference-of-convex (DC-constrained DC) program.
  • We solve the per-iteration program with sequential convex programming (SCP), the standard solver for difference-of-convex problems, and give convergence guarantees under mild conditions, bridging the difference-of-convex optimization and RL literatures.
  • Empirically, multi-step Convex-Concave RL (CCRL) wins on diagnostic MDPs where credit must propagate across a horizon (its advantage growing with the dependency length), is competitive with a tuned PPO on classic control, and on a realistic, stochastic, mid-horizon healthcare domain converges markedly faster than tuned PPO to the same near-optimal survival, with an 11.3% higher area under the training curve.

Sources (1)

  • [1]Convex-Concave Reinforcement Learning
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:59 PM
    Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today.
    Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even under a direct policy parameterization, and the field has largely responded by avoiding it: optimizing convex surrogate approximations of the return under trust-region constraints (NPG, TRPO, PPO, AWR).

Extractive summary: sentences quoted from the sources.