Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Published
Source
arXiv
Paper number
904
Field
AI / General
arXiv ID
2608.13417

Key points

  • Even when final scores are the same, the bottlenecks differ. GPT-5.5 is strong at execution, while Gemini-3.1-Pro is strong at feedback control, so their process profiles diverge.
  • The gap between models is much larger on average performance, avg@3 at 0.237, than on best@3 at 0.122. Reliability is the deciding factor.
  • Only 3 of the 252 top-seed solutions had real methodological novelty.
  • Experience transfer cuts both ways. DeepSeek-V4-Pro improves by +0.093, while Gemini-3.1-Pro degrades by -0.017.
  • The harness, or agent operating system, mainly affects execution stability rather than peak performance.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)