HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

Published
Source
arXiv
Paper number
1060
Field
Agents / Infrastructure
arXiv ID
2609.01437

Key points

  • HarnessDev defines a harness as execution infrastructure reused across tasks rather than a single response, separates creation from evolution, and evaluates both capability and execution token cost.
  • It used 6 generation models and 2,207 unique tasks across 4 domains and 5 benchmarks. Generated harnesses were strong in writing and machine-learning experiments but lagged behind human-built baselines in coding, search, and research.
  • All 18 coding harnesses implemented execution loops, but state and memory were their weakest areas. Although 11 declared a State class, only 1 provided a storage interface and only 1 provided periodic checkpoints, and no checkpoint event appeared in 26,679 execution traces.
  • Analysis of 73 official versions and 64 version transitions across 9 evolution runs found that visible and hidden scores moved in the same direction 53.1% of the time, and the final selected version was optimal on hidden tasks in only 2/9 cases.
  • In evolution experiments with Gemini fixed as the execution model, only the Opus family improved on hidden tasks, while the other 3 families worsened, showing that harness improvements can overfit a particular model and public feedback.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)