Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
- Published
- Source
- arXiv
- Paper number
- 904
- Field
- AI / General
- arXiv ID
- 2608.13417
Key points
- Even when final scores are the same, the bottlenecks differ. GPT-5.5 is strong at execution, while Gemini-3.1-Pro is strong at feedback control, so their process profiles diverge.
- The gap between models is much larger on average performance, avg@3 at 0.237, than on best@3 at 0.122. Reliability is the deciding factor.
- Only 3 of the 252 top-seed solutions had real methodological novelty.
- Experience transfer cuts both ways. DeepSeek-V4-Pro improves by +0.093, while Gemini-3.1-Pro degrades by -0.017.
- The harness, or agent operating system, mainly affects execution stability rather than peak performance.
Paper links
External research summaries. These are not HDATF publications or measured product results.