ResearchResearch paperReinforcement Learning1 source · Oct 6, 2026

Pareto-Optimal Entropy-Regularized Trajectory Optimization

We introduce Pareto-Optimal Entropy-Regularized DDP (PER-DDP), an entropy-regularized population framework derived from the free-energy/relative-entropy inequality.

Key points

  • Trajectory optimization (TO) under nonlinear dynamics, actuation limits and collision avoidance constraints is a fundamental problem in robotics, albeit especially challenging due to its highly non-convex nature.
  • For this setting, Differential Dynamic Programming (DDP) is an efficient second-order shooting method, yet its local structure renders it vulnerable to suboptimal basins.
  • Our method combines prior-guided sampling that shapes exploration around each retained trajectory, with expanded rollout evaluations that probe these sampling policies beyond the few retained candidates, and Pareto filtering for preserving task-constraint alternatives across iterations.
  • This decouples sampling effort from the optimization population size and broadens exploration without sacrificing the second-order structure that makes DDP effective.

Sources (1)

  • [1]Pareto-Optimal Entropy-Regularized Trajectory Optimization
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:25 PM
    We introduce Pareto-Optimal Entropy-Regularized DDP (PER-DDP), an entropy-regularized population framework derived from the free-energy/relative-entropy inequality.
    Trajectory optimization (TO) under nonlinear dynamics, actuation limits and collision avoidance constraints is a fundamental problem in robotics, albeit especially challenging due to its highly non-convex nature.

Extractive summary: sentences quoted from the sources.

Related