VOMMI: Collecting and Leveraging Portable Demonstrations for Mobile Manipulation
We present the Visual-Odometry-Conditioned Mobile Manipulation Interface (VOMMI), a portable demonstration collection and learning framework that connects portable RGB demonstrations to vision-language-action (VLA) post-training through offline trajectory reconstruction and online visual-motion conditioning.
ProofPaper ↗
Key points
- Portable mobile-manipulation demonstrations can help alleviate data scarcity for embodied intelligence, but obtaining reliable, low-cost, and robot-free motion supervision from RGB observations remains challenging.
- Existing approaches often rely on teleoperation or specialized devices equipped with additional sensing hardware, while directly using estimated visual odometry (VO) trajectories can introduce inconsistencies due to accumulated drift and imperfect motion supervision.
- VOMMI synchronizes body and hand views to capture navigation context and local object interactions without requiring human-robot kinematic correspondence calibration.
- R2-VO refines offline demonstration trajectories using sparse geometric anchors and produces causal local-motion tokens over multiple prediction horizons for online policy conditioning.
Sources (1)
- [1]VOMMI: Collecting and Leveraging Portable Demonstrations for Mobile ManipulationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 12:10 PM
We present the Visual-Odometry-Conditioned Mobile Manipulation Interface (VOMMI), a portable demonstration collection and learning framework that connects portable RGB demonstrations to vision-language-action (VLA) post-training through offline trajectory reconstruction and online visual-motion conditioning.
Portable mobile-manipulation demonstrations can help alleviate data scarcity for embodied intelligence, but obtaining reliable, low-cost, and robot-free motion supervision from RGB observations remains challenging.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Adapting Vision-Language-Action Models to Unknown Visual Disruptions During Execution
- Oct 6, 2026StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action Models
- Oct 6, 2026CARE: Certifying Acceleration for Vision-Language-Action Inference
- Oct 2, 2026FastOPD: On-Policy Distillation for Lightweight VLA Deployment
- Jul 30, 2026Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
- Jul 28, 2026Gemini Robotics 2 brings whole body intelligence to robots