StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action Models
Based on this insight, we introduce StairVLA, a stage-aware hierarchical action generation framework that uses partially denoised actions as a natural interface between coarse long-horizon action generation and local refinement.
ProofPaper ↗
Key points
- Vision-language-action (VLA) models increasingly rely on diffusion- or flow-matching-based action heads to generate continuous robot actions.
- However, we observe that the conditioning focus naturally shifts across denoising stages: early stages combine language instructions and visual observations to establish a coarse action trajectory, whereas later stages place greater emphasis on current visual observations for action alignment.
- A high-level VLA performs early denoising to produce a reusable long-horizon partially denoised action trajectory, while a lightweight refiner operates at a higher frequency to refine local action chunks using the latest observations.
- On LIBERO, our GR00T-style instantiation improves average success from 96.5% to 97.8% while reducing amortized inference latency from 115.0 ms to 44.2 ms per action chunk.
Sources (1)
- [1]StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:53 AM
Based on this insight, we introduce StairVLA, a stage-aware hierarchical action generation framework that uses partially denoised actions as a natural interface between coarse long-horizon action generation and local refinement.
Vision-language-action (VLA) models increasingly rely on diffusion- or flow-matching-based action heads to generate continuous robot actions.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026ESP: Energy-Score Policy for One-Step Multimodal Action Generation
- Oct 6, 2026SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining
- Oct 6, 2026CARE: Certifying Acceleration for Vision-Language-Action Inference
- Oct 2, 2026FastOPD: On-Policy Distillation for Lightweight VLA Deployment
- Jul 30, 2026Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
- Jul 28, 2026Gemini Robotics 2 brings whole body intelligence to robots