BrickBench: Evaluating Agentic Brick Design
Peter Kulits, Yiqing Xu, R. Kenny Jones, Cordelia Schmid, Jiajun Wu
Abstract
We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at http://www.brickben.ch
Create a lesson
Related papers
On the estimation and validity of AI time horizons---a statistical look at the METR plot
Drew T. Nguyen, William Fithian
Ecology of AI Agents: Collaboration Creates a Population Threshold for Takeoff
Erin Crawley, Hidenori Tanaka
Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark
Christopher M. Stewart, Preston Botter, Natalie Sarabosing et al.
HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling
Qun Dai, Liangjian Wen, Jiang Duan et al.
GeoReform: Reflective Formalization Evolution for Multimodal Geometry Problem Solving
Jialu Wang, Ruichen Zhang, Xiaoou Liu et al.
OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport
Babak Barazandeh, Connor Swanson, Chinmay Kulkarni et al.