ResearchResearch paperRobotics & Embodied AI1 source · Oct 6, 2026

AssemState: Manual and Physical-State-Guided Reasoning for Zero-shot Furniture Assembly

To study this problem, we propose AssemState, a zero-shot framework for manual and physical-state-guided furniture assembly.

Key points

  • Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult.
  • Furniture assembly requires not only recovering step-level operations from diagrammatic manuals, but also translating semantic attachment relations into 6D pose updates that enable parts to physically interact with the environment and previously assembled components.
  • It firstly employs anchor-guided boundary assembly states to decompose manual pages into single-part operations and recover an assembly-tree.
  • These results show that AssemState improves operation-structure recovery and selected local pose metrics, while MLLMs remain limited for spatial relationship reasoning.

Sources (1)

  • [1]AssemState: Manual and Physical-State-Guided Reasoning for Zero-shot Furniture Assembly
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 02:35 PM
    To study this problem, we propose AssemState, a zero-shot framework for manual and physical-state-guided furniture assembly.
    Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult.

Extractive summary: sentences quoted from the sources.

Related