TempoBridge: Language-Guided Tempo Control for Vision-Language-Action Policies
We introduce TempoBridge, a lightweight framework that uses frozen VLA representations to modulate actions according to tempo cues in the instruction at each task phase, without additional tempo-conditioned robot demonstrations or tempo-specific base-policy fine-tuning.
ProofPaper ↗
Key points
- Vision-Language-Action (VLA) models are effective at understanding what task to perform, but provide limited control over how it should be executed, such as moving quickly or slowly.
- TempoBridge extracts tempo cues from contextual VLM representations, aligns them with task progress through a causal phase router, and modulates nominal motion commands during execution.
- It also preserves near-baseline performance when no tempo cue is present and generalizes to unseen tempo expressions without additional training.
- Experiments on a physical robot further demonstrate language-conditioned tempo modulation in real-world manipulation.
Sources (1)
- [1]TempoBridge: Language-Guided Tempo Control for Vision-Language-Action PoliciesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:07 AM
We introduce TempoBridge, a lightweight framework that uses frozen VLA representations to modulate actions according to tempo cues in the instruction at each task phase, without additional tempo-conditioned robot demonstrations or tempo-specific base-policy fine-tuning.
Vision-Language-Action (VLA) models are effective at understanding what task to perform, but provide limited control over how it should be executed, such as moving quickly or slowly.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Q-Learning with Scalar Adjoint Matching
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026Co-Evolving Robot Orchestrators and Policies through Deployment
- Oct 6, 2026VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models
- Jul 30, 2026Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
- Jun 10, 2026DiffusionGemma: 4x faster text generation