SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

Published
Source
arXiv
Paper number
992
Field
LLMs / NLP
arXiv ID
2608.23564

Key points

  • Behavioral tests alone cannot filter out submissions that leave the original implementation intact, so when agents are tasked with replacing a repository's entire stack, whether the replacement actually occurred must be checked separately.
  • It created tasks from 20 real open-source repositories covering 4 types of technical debt: languages, frameworks, platforms, and build tools.
  • It scores in 3 stages using migration audits, 130,118 fixed behavioral tests, and adversarial validation by 6 coding agents. However, the paper draws a clear boundary: the absence of counterexamples is not proof of behavioral equivalence, but only the strongest evidence this benchmark can provide.
  • Only 5.4% (28) of 520 runs passed every stage, and no model solved 13 of the 20 tasks.
  • Capabilities varied widely: build-tool replacement achieved some success (31.4 points), while language replacement almost always failed (5.6 points).

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)