REACT: Rolling Denoising and Dual Decoupling for Reactive Robot Control with VLA Models
Flow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities.
Key points
- We introduce REACT, a rolling-denoising framework that makes flow-based VLAs more reactive while preserving long-horizon context.
- Instead of regenerating entire action chunks from scratch, REACT maintains a persistent action buffer with staggered flow timesteps.
- To support real-time control, we further introduce dual decoupling, which separates sensing, VLM encoding, DiT denoising, and action execution, enabling high-frequency observation updates and action streaming under practical compute constraints.
- Across the RoboTwin 2.0 simulation benchmark and real-world tasks spanning bimanual manipulation and dynamic control on multiple robot platforms, REACT improves task success and reduces reaction latency while producing smoother trajectories than frequent-replanning and asynchronous baselines.
Sources (1)
- [1]REACT: Rolling Denoising and Dual Decoupling for Reactive Robot Control with VLA ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:09 PM
Flow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities.
We introduce REACT, a rolling-denoising framework that makes flow-based VLAs more reactive while preserving long-horizon context.
Extractive summary: sentences quoted from the sources.