CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Published
Source
arXiv
Paper number
1104
Field
AI / General
arXiv ID
2609.22068

Key points

  • The core idea is building training environments from source code alone, without development artifacts like issues or commits.
  • Agents explore the code themselves to write the task, then auto-grade solutions with tests grounded in the original code's execution.
  • It extracted 5,545 tasks from 3,185 open-source projects, covering 23 programming languages and 15 technical domains.
  • Training improved bug fixing (DeepSWE) by +11.7%p, whole-program construction (ProgramBench) by +17%p, and terminal work (Terminal-Bench v2.1) by +8.5%p.
  • The filtered 5,000 high-quality tasks beat an unrefined 8,000-task set. Quality won over quantity.
  • As training progressed, agents explored the codebase more and self-verified more often.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)