ResearchResearch paperRobotics & Embodied AI · Reinforcement Learning · Training & Scaling1 source · Oct 8, 2026

A Geometric Approach to Soft Actor-Critic with Zonotopes for Locomotion Learning

We propose GeZo-SAC, which uses auxiliary geometric representations to adapt critic pessimism to the state and action.

Key points

  • Off-policy actor--critic methods control overestimation bias by taking the minimum of two critics.
  • Alongside its scalar value, each critic predicts a set of generators defining a zonotope.
  • At inference, the deployed policy is an unmodified SAC actor, since the generators are used only on the critic side during training.Across four MuJoCo-v5 locomotion benchmarks and six off-policy baselines, GeZo-SAC achieves the highest mean return on Ant-v5 and Hopper-v5 and remains competitive with other methods on the remaining tasks.
  • Our analysis further shows that GeZo-SAC achieves the lowest average actuator work and action effort per metre among the evaluated methods, while maintaining near-zero measured overestimation frequency across all four environments.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related