Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver

Published
Source
arXiv
Paper number
173
Field
Agents / Coding / Evaluation
arXiv ID
2604.25067

Key points

  • Existing AI capability benchmarks may fail to provide sufficiently early warning signs of recursive self-improvement (RSI).
  • There is a need to measure AI's ability to autonomously implement end-to-end machine learning pipelines based on established AI research results.
  • Evaluation of advanced AI systems is complicated by potential strategic behavior such as sandbagging or deliberate performance degradation.
  • We developed a benchmark that tasks frontier coding agents with autonomously implementing an AlphaZero-style self-play ML pipeline for Connect Four within a 3-hour time limit on consumer hardware.
  • We evaluated four frontier coding agents, Gemini 3.1 Pro, Claude Opus 4.6, Claude Opus 4.7, and GPT-5.4, in an isolated sandbox Docker environment.
  • We investigated anomalous time-budget usage patterns by running sandbagging probes on GPT-5.4 with different prompt strategies, such as Hobbyist versus RSI Alert, and different execution environments.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)