SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution

Published
Source
arXiv
Paper number
955
Field
Software Engineering
arXiv ID
2608.18933

Key points

  • SkillForge takes a proactive approach: it synthesizes issues directly from the repository and accumulates knowledge in advance, without relying on issue history or incurring costly per-issue test-time exploration.
  • The synthetic issues are test-covered bugs that actually execute, so the learning signal is verified through program behavior.
  • The distilled knowledge is organized in a two-tier store of global diagnostic skills and local intervention skills, then used through global initialization and just-in-time injection during issue resolution.
  • On SWE-bench Verified, SkillForge reaches 72.2% Pass@1 with DeepSeek-V3.2, an improvement of 5.8 percentage points, and 60.6% with GPT-5-mini, an improvement of 5.6 percentage points. On SWE-bench Pro, it improves the two models by 5.8 and 4.1 percentage points, respectively.
  • The type distribution of the synthetic issues does not overlap with that of the real evaluation issues. This indicates that the method targets new model-induced vulnerabilities rather than recycling the evaluation set.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)