Residual Modeling Closes the Regression and Generative Policy Gap in Robot Learning
We revisit this gap from the perspective of statistical modeling: how action-prediction residuals shape policy optimization.
ProofPaper ↗
Key points
- Learning from demonstration has enabled impressive robot behaviors.
- A common choice for policy learning is to use diffusion or flow matching (Flow-Policies), which often outperforms direct action regression trained with mean squared error (MSE-Policies).
- Motivated by these findings, we introduce heteroscedastic Student-t action regression (HT-Policies), which learns input-dependent residual scales and reduces the influence of heavy tails.
- Across four simulation benchmarks and real-robot evaluations, HT-Policies achieves success rates competitive with generative policy baselines, both when trained from scratch and from pretrained vision-language-action and world-action models, despite being faster in training and inference.
Sources (1)
- [1]Residual Modeling Closes the Regression and Generative Policy Gap in Robot LearningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:16 PM
We revisit this gap from the perspective of statistical modeling: how action-prediction residuals shape policy optimization.
Learning from demonstration has enabled impressive robot behaviors.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
- Oct 8, 2026Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
- Oct 8, 2026Normalizing Trajectory Models
- Oct 7, 2026Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
- Oct 7, 2026Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding
- Oct 7, 2026Q-Learning with Scalar Adjoint Matching