SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

Published
Source
arXiv
Paper number
859
Field
LLMs / NLP
arXiv ID
2608.09802

Key points

  • It quantifies the problem that 60% of existing benchmarks contain bad tests that are either too narrow or too broad.
  • It builds a multilingual refactoring benchmark across seven programming languages: Python, Java, TypeScript, Go, C, C++, and Rust.
  • The tasks are large, with an average of 11.4 files and 261.6 lines of code to change, which makes them more than three times more complex than prior benchmarks.
  • GLM-5 shows strong value for money by matching Claude Sonnet 4.6 at about one twentieth of the cost.
  • Failure analysis shows that the main problem is not modifying enough of the files that actually need to change.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)