BrickBench: Evaluating Agentic Brick Design
- Published
- Source
- arXiv
- Paper number
- 1184
- Field
- AI Agents
- arXiv ID
- 2610.12452
Key points
- The best coding agent satisfied about 95% of prompt semantic requirements and produced physically valid assemblies for nearly every prompt.
- Removing the BrickAgent tool environment crashed valid-assembly rates from 100% to 40% for GPT-6 Astra and to under 1% for GPT-5.6 Luna, showing that grounded validation tools are the key to success.
- General coding agents outperformed specialized LEGO-generation models (BrickNet-14B, BrickGPT), though at 600 to 6800 times the cost.
- Human raters chose the human design in 323 of 360 comparisons (about 90%), showing that meeting requirements and designing well are different problems.
- They computed the design-quality ELO metric with a VLM judge and validated its ranking agreement with human preferences (Kendall tau = 0.89).
Paper links
External research summaries. These are not HDATF publications or measured product results.