SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

Published
Source
arXiv
Paper number
963
Field
LLMs / NLP
arXiv ID
2608.19799

Key points

  • It assembled 119 tasks from 98 repositories across 20 scientific fields, including chemistry, materials science, and biology, to evaluate scientific-software repair capabilities.
  • Even the best agent achieved only a 47.9% single-attempt success rate, with a large gap between the public-test score of 96.6% and the full success rate.
  • It classified failures into four types: insufficient scientific knowledge or abstraction, superficial repairs, incomplete integration, and failure to generalize to new cases.
  • In one case, adding scientific information reduced GPT-5.6's success rate from 36.3% to 31.9%, showing that knowledge must be validated alongside execution evidence.
  • The number of tasks per field remains small, and analysis of how scientific knowledge is used in actual repairs is still preliminary, so cross-field rankings and generalization require cautious interpretation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)