Evo-Bench: Can Language Models Improve Agent Harness?

Published
Source
arXiv
Paper number
875
Field
LLMs / NLP
arXiv ID
2608.09096

Key points

  • It creates the first benchmark for evaluating whether an LLM can autonomously improve the agent harness, meaning the execution code.
  • It introduces a strict design that avoids overfitting through harness-sensitivity-based task selection and separate validation and evaluation splits.
  • Evaluation on nine frontier models shows gains of up to 16.6 points and comes close to the human-designed baseline.
  • In the general domain, autonomous evolution beats the human design, but it shows limits in office workflows.
  • It confirms that the evolved harness acts as a transferable reasoning structure for other policy models.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)