SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
- Published
- Source
- arXiv
- Paper number
- 963
- Field
- LLMs / NLP
- arXiv ID
- 2608.19799
Key points
- It assembled 119 tasks from 98 repositories across 20 scientific fields, including chemistry, materials science, and biology, to evaluate scientific-software repair capabilities.
- Even the best agent achieved only a 47.9% single-attempt success rate, with a large gap between the public-test score of 96.6% and the full success rate.
- It classified failures into four types: insufficient scientific knowledge or abstraction, superficial repairs, incomplete integration, and failure to generalize to new cases.
- In one case, adding scientific information reduced GPT-5.6's success rate from 36.3% to 31.9%, showing that knowledge must be validated alongside execution evidence.
- The number of tasks per field remains small, and analysis of how scientific knowledge is used in actual repairs is still preliminary, so cross-field rankings and generalization require cautious interpretation.
Paper links
External research summaries. These are not HDATF publications or measured product results.