Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
- Published
- Source
- arXiv
- Paper number
- 625
- Field
- AI / General
- arXiv ID
- 2607.12463
Key points
- It designed function-aware FIM mid-training by exploiting the similarity between a coding agent's action-observation-continuation structure and the function call-return structure.
- It selected functions to mask using program-dependence graphs and criteria for complexity and inferability, then trained on 2.6 billion tokens from 968 GitHub repositories.
- SWE-Bench-Verified improved by 2.8, 3.0, and 3.2 points on Qwen-family 7B, 14B, and 8B models, respectively.
- The effect persisted across multiple agent post-training procedures and also mitigated capability degradation in general coding and non-coding tool use.
- The training corpus and evaluations focused on Python, and generating reasoning explanations depended on a teacher model, so other languages and teacher-free settings remain unverified.
Paper links
External research summaries. These are not HDATF publications or measured product results.