BrickBench: Evaluating Agentic Brick Design
We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design.
Key points
- Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built.
- To do so, it must select parts from a discrete library and reason jointly about local and global constraints.
- We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs.
- We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs.
Sources (2)
- [1]BrickBench: Evaluating Agentic Brick DesignHugging Face Daily Papers · Oct 8, 12:00 AM
We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design.
Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built.
- [2]BrickBench: Evaluating Agentic Brick DesignarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:58 PM · same content
Extractive summary: sentences quoted from the sources.