The Verification Horizon: No Silver Bullet for Coding Agent Rewards

Published
Source
arXiv
Paper number
511
Field
AI / General
arXiv ID
2606.26300

Key points

  • When quality judging and trajectory monitoring are added to unit-test rewards, the SWE-Bench hacking solve rate drops sharply from 28.57% to 0.56%, while the normal solve rate rises from 40.22% to 60.53%.
  • For frontend tasks, introducing a browser-based interaction evaluator defends against length attacks that are vulnerable to static code inspection.
  • User-feedback rewards achieve up to a 13.3-point improvement across five internal coding-agent benchmarks.
  • For autonomous-evaluator data filtered for long-horizon tasks, RFT gains +1.91 points over random sampling, from 21.61 to 23.52.
  • There is no single mechanism that satisfies all three dimensions of verification, scalability, fidelity, and robustness, so task-specific design is essential.
  • Verification is not auxiliary but core training infrastructure, and it must co-evolve continuously with policy capability growth.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)