AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
- Published
- Source
- arXiv
- Paper number
- 972
- Field
- AI / General
- arXiv ID
- 2608.20318
Key points
- It created a benchmark that freezes 10 research repositories covering 10 training-algorithm families, lets an agent modify the training code for 4 hours, then retrains from scratch for up to 12 hours and scores the result with a hidden evaluator.
- The average score across 29 configurations of 6 systems was 0.166, and even the best score of 0.250 closed less than one-fifth of the gap between the original algorithm (0.1) and the optimum (1.0).
- Of 263 submissions with changes, 141 did not modify the training procedure at all. The 122 that did scored an average of 0.226, substantially outperforming the rest (0.126).
- Increasing reasoning effort raised the proportion of attempts that addressed the algorithmic layer from 8% to 64%, and the average score from 0.094 to 0.196.
- It released the task suite, evaluators, and all scored submissions so that evolving systems can be measured again using the same yardstick.
Paper links
External research summaries. These are not HDATF publications or measured product results.