AION
Research paperRobotics & Embodied AI1 source · Oct 8, 2026

ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks

Long-horizon mobile manipulation requires a robot to navigate multi-room environments and execute a sequence of manipulation skills under a single natural language instruction.

Key points

  • Learning and evaluating these skills present three challenges: similar observations under a fixed task instruction may make skill selection ambiguous; even when a preceding skill succeeds, the robot state inherited by the next skill may deviate from its demonstrated starting states and affect execution; and task-level metrics hinder skill-specific diagnosis, while early failures leave later skills untested.
  • We therefore introduce ManiUnit, a manipulation skill dataset and benchmark built from 50 BEHAVIOR-1K activities.
  • Correspondingly, ManiUnit pairs each segment with an explicit subtask instruction; measures sensitivity to perturbations of the robot's starting base position or joint configuration; and restores intermediate simulator states and defines local success conditions so that each skill can be evaluated without executing preceding stages.
  • On two long-horizon activities, a skill policy trained on ManiUnit segments achieves 78.7% local manipulation success, compared with 49.3% for a task policy trained on complete demonstrations.

Sources (1)

  • [1]ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:00 PM
    Long-horizon mobile manipulation requires a robot to navigate multi-room environments and execute a sequence of manipulation skills under a single natural language instruction.
    Learning and evaluating these skills present three challenges: similar observations under a fixed task instruction may make skill selection ambiguous; even when a preceding skill succeeds, the robot state inherited by the next skill may deviate from its demonstrated starting states and affect execution; and task-level metrics hinder skill-specific diagnosis, while early failures leave later skills

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
  2. Oct 8, 2026Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
  3. Oct 7, 2026Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
  4. Oct 7, 2026Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding
  5. Oct 7, 2026Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
  6. Oct 7, 2026Q-Learning with Scalar Adjoint Matching

Related