CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
- Published
- Source
- arXiv
- Paper number
- 1104
- Field
- AI / General
- arXiv ID
- 2609.22068
Key points
- The core idea is building training environments from source code alone, without development artifacts like issues or commits.
- Agents explore the code themselves to write the task, then auto-grade solutions with tests grounded in the original code's execution.
- It extracted 5,545 tasks from 3,185 open-source projects, covering 23 programming languages and 15 technical domains.
- Training improved bug fixing (DeepSWE) by +11.7%p, whole-program construction (ProgramBench) by +17%p, and terminal work (Terminal-Bench v2.1) by +8.5%p.
- The filtered 5,000 high-quality tasks beat an unrefined 8,000-task set. Quality won over quantity.
- As training progressed, agents explored the codebase more and self-verified more often.
Paper links
External research summaries. These are not HDATF publications or measured product results.