Evo-Bench: Can Language Models Improve Agent Harness?
- Published
- Source
- arXiv
- Paper number
- 875
- Field
- LLMs / NLP
- arXiv ID
- 2608.09096
Key points
- It creates the first benchmark for evaluating whether an LLM can autonomously improve the agent harness, meaning the execution code.
- It introduces a strict design that avoids overfitting through harness-sensitivity-based task selection and separate validation and evaluation splits.
- Evaluation on nine frontier models shows gains of up to 16.6 points and comes close to the human-designed baseline.
- In the general domain, autonomous evolution beats the human design, but it shows limits in office workflows.
- It confirms that the evolved harness acts as a transferable reasoning structure for other policy models.
Paper links
External research summaries. These are not HDATF publications or measured product results.