PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies
Vision-Language-Action (VLA) policies typically predict actions at fixed time intervals, coupling the route a robot follows with its execution pace.
Key points
- Our key insight is to bring the path-time parameterization of classical motion planning into the learned action representation of a VLA.
- We introduce PathTime-VLA, which represents motion as a progress-indexed interaction path $X(s)$ and a positive interval-time profile.
- This representation supports a staged post-training procedure: demonstrations and DAgger interventions establish a target-domain prior, Speed-DQN learns execution multipliers from robot interaction, and Path-AWR uses rollout outcomes to refine the diffusion path generator.
- A path-conditioned action expert realizes the resulting motions while maintaining distinct learning interfaces for path generation and execution timing.
Sources (1)
- [1]PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action PoliciesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 11:53 AM
Vision-Language-Action (VLA) policies typically predict actions at fixed time intervals, coupling the route a robot follows with its execution pace.
Our key insight is to bring the path-time parameterization of classical motion planning into the learned action representation of a VLA.
Extractive summary: sentences quoted from the sources.