Pareto-Optimal Entropy-Regularized Trajectory Optimization
We introduce Pareto-Optimal Entropy-Regularized DDP (PER-DDP), an entropy-regularized population framework derived from the free-energy/relative-entropy inequality.
ProofPaper ↗
Key points
- Trajectory optimization (TO) under nonlinear dynamics, actuation limits and collision avoidance constraints is a fundamental problem in robotics, albeit especially challenging due to its highly non-convex nature.
- For this setting, Differential Dynamic Programming (DDP) is an efficient second-order shooting method, yet its local structure renders it vulnerable to suboptimal basins.
- Our method combines prior-guided sampling that shapes exploration around each retained trajectory, with expanded rollout evaluations that probe these sampling policies beyond the few retained candidates, and Pareto filtering for preserving task-constraint alternatives across iterations.
- This decouples sampling effort from the optimization population size and broadens exploration without sacrificing the second-order structure that makes DDP effective.
Sources (1)
- [1]Pareto-Optimal Entropy-Regularized Trajectory OptimizationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:25 PM
We introduce Pareto-Optimal Entropy-Regularized DDP (PER-DDP), an entropy-regularized population framework derived from the free-energy/relative-entropy inequality.
Trajectory optimization (TO) under nonlinear dynamics, actuation limits and collision avoidance constraints is a fundamental problem in robotics, albeit especially challenging due to its highly non-convex nature.
Extractive summary: sentences quoted from the sources.