SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios

Published
Source
arXiv
Paper number
104
Field
Benchmarks / Code
arXiv ID
2512.18470

Key points

  • Existing AI coding-agent benchmarks mostly focus on isolated bug fixes or single-issue feature implementations, as seen in SWE-Bench.
  • These benchmarks do not adequately capture the long-horizon nature of real software engineering, which involves continuous evolution, multi-file coordination, and iterative updates to existing codebases.
  • A large share of software development is spent maintaining and evolving legacy code, which requires persistent understanding, adaptive planning, and holistic system changes that current methods do not evaluate.
  • SWE-EVO presents a benchmark of 48 high-quality software evolution tasks derived from real release notes, requiring the agent to switch between two distinct versions of a codebase.
  • The agent must interpret a high-level software requirements specification, or SRS, and implement multi-step, multi-file changes across the entire codebase while preserving functionality and ensuring that new tests pass.
  • Performance is evaluated with solve rate under strict pass or fail rules, patch apply rate, and edit rate as a finer-grained measure of progress, and it is supplemented with LLM-as-a-judge analysis to classify failure modes.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)