GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

Published
Source
arXiv
Paper number
438
Field
LLMs / NLP
arXiv ID
2606.17861

Key points

  • It is the first work to define three requirements for engine-based game generation evaluation: Engine Grounding, Artifact Completeness, and Interactive Verification.
  • It introduces a real engine benchmark with 140 Godot tasks across 15 game genres.
  • It uses a three-stage interaction-grounded evaluation pipeline that goes from headless build to play video to rubric-based multimodal judging.
  • The strongest agent, Opus 4.7 high, reaches 41.46%, while most agents remain below 40%, which shows that end-to-end game generation is still unsolved.
  • Agents can implement mechanics, but they fall far short on content depth, visual feedback, and presentation polish.
  • Mechanics, content, and visuals are moderately correlated at r=0.5 to 0.6, but art and presentation are weakly coupled at r=0.11, which shows that the capability splits only partially.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)