A Geometric Approach to Soft Actor-Critic with Zonotopes for Locomotion Learning
We propose GeZo-SAC, which uses auxiliary geometric representations to adapt critic pessimism to the state and action.
ProofPaper ↗
Key points
- Off-policy actor--critic methods control overestimation bias by taking the minimum of two critics.
- Alongside its scalar value, each critic predicts a set of generators defining a zonotope.
- At inference, the deployed policy is an unmodified SAC actor, since the generators are used only on the critic side during training.Across four MuJoCo-v5 locomotion benchmarks and six off-policy baselines, GeZo-SAC achieves the highest mean return on Ant-v5 and Hopper-v5 and remains competitive with other methods on the remaining tasks.
- Our analysis further shows that GeZo-SAC achieves the lowest average actuator work and action effort per metre among the evaluated methods, while maintaining near-zero measured overestimation frequency across all four environments.
Sources (1)
- [1]A Geometric Approach to Soft Actor-Critic with Zonotopes for Locomotion LearningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:11 PM
We propose GeZo-SAC, which uses auxiliary geometric representations to adapt critic pessimism to the state and action.
Off-policy actor--critic methods control overestimation bias by taking the minimum of two critics.
Extractive summary: sentences quoted from the sources.