SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution
- Published
- Source
- arXiv
- Paper number
- 955
- Field
- Software Engineering
- arXiv ID
- 2608.18933
Key points
- SkillForge takes a proactive approach: it synthesizes issues directly from the repository and accumulates knowledge in advance, without relying on issue history or incurring costly per-issue test-time exploration.
- The synthetic issues are test-covered bugs that actually execute, so the learning signal is verified through program behavior.
- The distilled knowledge is organized in a two-tier store of global diagnostic skills and local intervention skills, then used through global initialization and just-in-time injection during issue resolution.
- On SWE-bench Verified, SkillForge reaches 72.2% Pass@1 with DeepSeek-V3.2, an improvement of 5.8 percentage points, and 60.6% with GPT-5-mini, an improvement of 5.6 percentage points. On SWE-bench Pro, it improves the two models by 5.8 and 4.1 percentage points, respectively.
- The type distribution of the synthetic issues does not overlap with that of the real evaluation issues. This indicates that the method targets new model-induced vulnerabilities rather than recycling the evaluation set.
Paper links
External research summaries. These are not HDATF publications or measured product results.