BrickBench: Evaluating Agentic Brick Design

Published
Source
arXiv
Paper number
1184
Field
AI Agents
arXiv ID
2610.12452

Key points

  • The best coding agent satisfied about 95% of prompt semantic requirements and produced physically valid assemblies for nearly every prompt.
  • Removing the BrickAgent tool environment crashed valid-assembly rates from 100% to 40% for GPT-6 Astra and to under 1% for GPT-5.6 Luna, showing that grounded validation tools are the key to success.
  • General coding agents outperformed specialized LEGO-generation models (BrickNet-14B, BrickGPT), though at 600 to 6800 times the cost.
  • Human raters chose the human design in 323 of 360 comparisons (about 90%), showing that meeting requirements and designing well are different problems.
  • They computed the design-quality ELO metric with a VLM judge and validated its ranking agreement with human preferences (Kendall tau = 0.89).

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)