SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
- Published
- Source
- arXiv
- Paper number
- 104
- Field
- Benchmarks / Code
- arXiv ID
- 2512.18470
Key points
- Existing AI coding-agent benchmarks mostly focus on isolated bug fixes or single-issue feature implementations, as seen in SWE-Bench.
- These benchmarks do not adequately capture the long-horizon nature of real software engineering, which involves continuous evolution, multi-file coordination, and iterative updates to existing codebases.
- A large share of software development is spent maintaining and evolving legacy code, which requires persistent understanding, adaptive planning, and holistic system changes that current methods do not evaluate.
- SWE-EVO presents a benchmark of 48 high-quality software evolution tasks derived from real release notes, requiring the agent to switch between two distinct versions of a codebase.
- The agent must interpret a high-level software requirements specification, or SRS, and implement multi-step, multi-file changes across the entire codebase while preserving functionality and ensuring that new tests pass.
- Performance is evaluated with solve rate under strict pass or fail rules, patch apply rate, and edit rate as a finer-grained measure of progress, and it is supplemented with LLM-as-a-judge analysis to classify failure modes.
Paper links
External research summaries. These are not HDATF publications or measured product results.