AssemState: Manual and Physical-State-Guided Reasoning for Zero-shot Furniture Assembly
To study this problem, we propose AssemState, a zero-shot framework for manual and physical-state-guided furniture assembly.
ProofPaper ↗
Key points
- Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult.
- Furniture assembly requires not only recovering step-level operations from diagrammatic manuals, but also translating semantic attachment relations into 6D pose updates that enable parts to physically interact with the environment and previously assembled components.
- It firstly employs anchor-guided boundary assembly states to decompose manual pages into single-part operations and recover an assembly-tree.
- These results show that AssemState improves operation-structure recovery and selected local pose metrics, while MLLMs remain limited for spatial relationship reasoning.
Sources (1)
- [1]AssemState: Manual and Physical-State-Guided Reasoning for Zero-shot Furniture AssemblyarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 02:35 PM
To study this problem, we propose AssemState, a zero-shot framework for manual and physical-state-guided furniture assembly.
Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult.
Extractive summary: sentences quoted from the sources.