VioLA: Learning Generalist Humanoid Control Policies from Human Data
We introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands.
ProofPaper ↗
Key points
- Teaching a humanoid to follow instructions with its whole body runs into two obstacles.
- Its action space is large and tightly coupled: legs, arms, and fingers must move together while the robot keeps its balance, which makes joint-level actions hard to learn.
- Their corresponding motion encoders map human and robot motion into the same latent spaces.
- A generalist policy trained on human demonstrations alone performs locomotion tasks on the real robot zero-shot.
Sources (1)
- [1]VioLA: Learning Generalist Humanoid Control Policies from Human DataarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:57 PM
We introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands.
Teaching a humanoid to follow instructions with its whole body runs into two obstacles.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026Understanding Jev, the new model everyone is talking about
- Oct 7, 2026Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
- Oct 7, 2026TempoBridge: Language-Guided Tempo Control for Vision-Language-Action Policies
- Oct 7, 2026Q-Learning with Scalar Adjoint Matching
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026Co-Evolving Robot Orchestrators and Policies through Deployment