HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- Published
- Source
- arXiv
- Paper number
- 1060
- Field
- Agents / Infrastructure
- arXiv ID
- 2609.01437
Key points
- HarnessDev defines a harness as execution infrastructure reused across tasks rather than a single response, separates creation from evolution, and evaluates both capability and execution token cost.
- It used 6 generation models and 2,207 unique tasks across 4 domains and 5 benchmarks. Generated harnesses were strong in writing and machine-learning experiments but lagged behind human-built baselines in coding, search, and research.
- All 18 coding harnesses implemented execution loops, but state and memory were their weakest areas. Although 11 declared a State class, only 1 provided a storage interface and only 1 provided periodic checkpoints, and no checkpoint event appeared in 26,679 execution traces.
- Analysis of 73 official versions and 64 version transitions across 9 evolution runs found that visible and hidden scores moved in the same direction 53.1% of the time, and the final selected version was optimal on hidden tasks in only 2/9 cases.
- In evolution experiments with Gemini fixed as the execution model, only the Opus family improved on hidden tasks, while the other 3 families worsened, showing that harness improvements can overfit a particular model and public feedback.
Paper links
External research summaries. These are not HDATF publications or measured product results.