ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks
Long-horizon mobile manipulation requires a robot to navigate multi-room environments and execute a sequence of manipulation skills under a single natural language instruction.
Key points
- Learning and evaluating these skills present three challenges: similar observations under a fixed task instruction may make skill selection ambiguous; even when a preceding skill succeeds, the robot state inherited by the next skill may deviate from its demonstrated starting states and affect execution; and task-level metrics hinder skill-specific diagnosis, while early failures leave later skills untested.
- We therefore introduce ManiUnit, a manipulation skill dataset and benchmark built from 50 BEHAVIOR-1K activities.
- Correspondingly, ManiUnit pairs each segment with an explicit subtask instruction; measures sensitivity to perturbations of the robot's starting base position or joint configuration; and restores intermediate simulator states and defines local success conditions so that each skill can be evaluated without executing preceding stages.
- On two long-horizon activities, a skill policy trained on ManiUnit segments achieves 78.7% local manipulation success, compared with 49.3% for a task policy trained on complete demonstrations.
Sources (1)
- [1]ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon TasksarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:00 PM
Long-horizon mobile manipulation requires a robot to navigate multi-room environments and execute a sequence of manipulation skills under a single natural language instruction.
Learning and evaluating these skills present three challenges: similar observations under a fixed task instruction may make skill selection ambiguous; even when a preceding skill succeeds, the robot state inherited by the next skill may deviate from its demonstrated starting states and affect execution; and task-level metrics hinder skill-specific diagnosis, while early failures leave later skills
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
- Oct 8, 2026Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
- Oct 7, 2026Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
- Oct 7, 2026Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding
- Oct 7, 2026Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
- Oct 7, 2026Q-Learning with Scalar Adjoint Matching