ResearchResearch paperRobotics & Embodied AI · Reinforcement Learning1 source · Oct 8, 2026

Residual Modeling Closes the Regression and Generative Policy Gap in Robot Learning

We revisit this gap from the perspective of statistical modeling: how action-prediction residuals shape policy optimization.

Key points

  • Learning from demonstration has enabled impressive robot behaviors.
  • A common choice for policy learning is to use diffusion or flow matching (Flow-Policies), which often outperforms direct action regression trained with mean squared error (MSE-Policies).
  • Motivated by these findings, we introduce heteroscedastic Student-t action regression (HT-Policies), which learns input-dependent residual scales and reduces the influence of heavy tails.
  • Across four simulation benchmarks and real-robot evaluations, HT-Policies achieves success rates competitive with generative policy baselines, both when trained from scratch and from pretrained vision-language-action and world-action models, despite being faster in training and inference.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
  2. Oct 8, 2026Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
  3. Oct 8, 2026Normalizing Trajectory Models
  4. Oct 7, 2026Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
  5. Oct 7, 2026Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding
  6. Oct 7, 2026Q-Learning with Scalar Adjoint Matching

Related