The Verification Horizon: No Silver Bullet for Coding Agent Rewards
- Published
- Source
- arXiv
- Paper number
- 511
- Field
- AI / General
- arXiv ID
- 2606.26300
Key points
- When quality judging and trajectory monitoring are added to unit-test rewards, the SWE-Bench hacking solve rate drops sharply from 28.57% to 0.56%, while the normal solve rate rises from 40.22% to 60.53%.
- For frontend tasks, introducing a browser-based interaction evaluator defends against length attacks that are vulnerable to static code inspection.
- User-feedback rewards achieve up to a 13.3-point improvement across five internal coding-agent benchmarks.
- For autonomous-evaluator data filtered for long-horizon tasks, RFT gains +1.91 points over random sampling, from 21.61 to 23.52.
- There is no single mechanism that satisfies all three dimensions of verification, scalability, fidelity, and robustness, so task-specific design is essential.
- Verification is not auxiliary but core training infrastructure, and it must co-evolve continuously with policy capability growth.
Paper links
External research summaries. These are not HDATF publications or measured product results.