GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
- Published
- Source
- arXiv
- Paper number
- 438
- Field
- LLMs / NLP
- arXiv ID
- 2606.17861
Key points
- It is the first work to define three requirements for engine-based game generation evaluation: Engine Grounding, Artifact Completeness, and Interactive Verification.
- It introduces a real engine benchmark with 140 Godot tasks across 15 game genres.
- It uses a three-stage interaction-grounded evaluation pipeline that goes from headless build to play video to rubric-based multimodal judging.
- The strongest agent, Opus 4.7 high, reaches 41.46%, while most agents remain below 40%, which shows that end-to-end game generation is still unsolved.
- Agents can implement mechanics, but they fall far short on content depth, visual feedback, and presentation polish.
- Mechanics, content, and visuals are moderately correlated at r=0.5 to 0.6, but art and presentation are weakly coupled at r=0.11, which shows that the capability splits only partially.
Paper links
External research summaries. These are not HDATF publications or measured product results.