SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
- Published
- Source
- arXiv
- Paper number
- 992
- Field
- LLMs / NLP
- arXiv ID
- 2608.23564
Key points
- Behavioral tests alone cannot filter out submissions that leave the original implementation intact, so when agents are tasked with replacing a repository's entire stack, whether the replacement actually occurred must be checked separately.
- It created tasks from 20 real open-source repositories covering 4 types of technical debt: languages, frameworks, platforms, and build tools.
- It scores in 3 stages using migration audits, 130,118 fixed behavioral tests, and adversarial validation by 6 coding agents. However, the paper draws a clear boundary: the absence of counterexamples is not proof of behavioral equivalence, but only the strongest evidence this benchmark can provide.
- Only 5.4% (28) of 520 runs passed every stage, and no model solved 13 of the 20 tasks.
- Capabilities varied widely: build-tool replacement achieved some success (31.4 points), while language replacement almost always failed (5.6 points).
Paper links
External research summaries. These are not HDATF publications or measured product results.