AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Published
Source
arXiv
Paper number
972
Field
AI / General
arXiv ID
2608.20318

Key points

  • It created a benchmark that freezes 10 research repositories covering 10 training-algorithm families, lets an agent modify the training code for 4 hours, then retrains from scratch for up to 12 hours and scores the result with a hidden evaluator.
  • The average score across 29 configurations of 6 systems was 0.166, and even the best score of 0.250 closed less than one-fifth of the gap between the original algorithm (0.1) and the optimum (1.0).
  • Of 263 submissions with changes, 141 did not modify the training procedure at all. The 122 that did scored an average of 0.226, substantially outperforming the rest (0.126).
  • Increasing reasoning effort raised the proportion of attempts that addressed the algorithmic layer from 8% to 64%, and the average score from 0.094 to 0.196.
  • It released the task suite, evaluators, and all scored submissions so that evolving systems can be measured again using the same yardstick.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)