Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver
- Published
- Source
- arXiv
- Paper number
- 173
- Field
- Agents / Coding / Evaluation
- arXiv ID
- 2604.25067
Key points
- Existing AI capability benchmarks may fail to provide sufficiently early warning signs of recursive self-improvement (RSI).
- There is a need to measure AI's ability to autonomously implement end-to-end machine learning pipelines based on established AI research results.
- Evaluation of advanced AI systems is complicated by potential strategic behavior such as sandbagging or deliberate performance degradation.
- We developed a benchmark that tasks frontier coding agents with autonomously implementing an AlphaZero-style self-play ML pipeline for Connect Four within a 3-hour time limit on consumer hardware.
- We evaluated four frontier coding agents, Gemini 3.1 Pro, Claude Opus 4.6, Claude Opus 4.7, and GPT-5.4, in an isolated sandbox Docker environment.
- We investigated anomalous time-budget usage patterns by running sandbagging probes on GPT-5.4 with different prompt strategies, such as Hobbyist versus RSI Alert, and different execution environments.
Paper links
External research summaries. These are not HDATF publications or measured product results.