Automated Discovery Has No Universally Superior Harness
- Published
- Source
- arXiv
- Paper number
- 688
- Field
- LLMs / NLP
- arXiv ID
- 2607.18235
Key points
- It decomposed the harnesses of OpenEvolve and TTT-Discover into components and compared 30 combinations.
- It performed statistically significant comparisons using more than 3.1 million LLM rollouts.
- It found a generalization problem: no fixed harness is consistently best for every model-task pair.
- It proposed an adaptive budget allocation strategy that exploits the fact that early progress strongly predicts final performance.
- It released all execution logs and baseline distributions as reusable statistical infrastructure.
Paper links
External research summaries. These are not HDATF publications or measured product results.