Automated Discovery Has No Universally Superior Harness

Published
Source
arXiv
Paper number
688
Field
LLMs / NLP
arXiv ID
2607.18235

Key points

  • It decomposed the harnesses of OpenEvolve and TTT-Discover into components and compared 30 combinations.
  • It performed statistically significant comparisons using more than 3.1 million LLM rollouts.
  • It found a generalization problem: no fixed harness is consistently best for every model-task pair.
  • It proposed an adaptive budget allocation strategy that exploits the fact that early progress strongly predicts final performance.
  • It released all execution logs and baseline distributions as reusable statistical infrastructure.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)