ResearchResearch paperRobotics & Embodied AI · Large Language Models1 source · Oct 7, 2026

TempoBridge: Language-Guided Tempo Control for Vision-Language-Action Policies

We introduce TempoBridge, a lightweight framework that uses frozen VLA representations to modulate actions according to tempo cues in the instruction at each task phase, without additional tempo-conditioned robot demonstrations or tempo-specific base-policy fine-tuning.

Key points

  • Vision-Language-Action (VLA) models are effective at understanding what task to perform, but provide limited control over how it should be executed, such as moving quickly or slowly.
  • TempoBridge extracts tempo cues from contextual VLM representations, aligns them with task progress through a causal phase router, and modulates nominal motion commands during execution.
  • It also preserves near-baseline performance when no tempo cue is present and generalizes to unseen tempo expressions without additional training.
  • Experiments on a physical robot further demonstrate language-conditioned tempo modulation in real-world manipulation.

Sources (1)

  • [1]TempoBridge: Language-Guided Tempo Control for Vision-Language-Action Policies
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:07 AM
    We introduce TempoBridge, a lightweight framework that uses frozen VLA representations to modulate actions according to tempo cues in the instruction at each task phase, without additional tempo-conditioned robot demonstrations or tempo-specific base-policy fine-tuning.
    Vision-Language-Action (VLA) models are effective at understanding what task to perform, but provide limited control over how it should be executed, such as moving quickly or slowly.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Q-Learning with Scalar Adjoint Matching
  2. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  3. Oct 6, 2026Co-Evolving Robot Orchestrators and Policies through Deployment
  4. Oct 6, 2026VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models
  5. Jul 30, 2026Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
  6. Jun 10, 2026DiffusionGemma: 4x faster text generation

Related