Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

Published
Source
arXiv
Paper number
719
Field
LLMs / NLP
arXiv ID
2607.22529

Key points

  • It uses reusable skills as the main carrier so that task diversity and verification reliability are both preserved.
  • A proposer, a solver, and a skill controller all evolve together in the reinforcement-learning loop.
  • The skill library creates about 20 new skills per iteration, updates existing skills, and drops useless ones.
  • With Qwen3-4B, it gains up to 42.9 points on tool use, BFCL, and 12.0 points on logical reasoning, ZebraLogic.
  • Dynamic skill evolution is the key, and it adds another 2.6 points in overall accuracy compared with self-learning without skills.
  • Even models that start out misaligned show a large turnaround effect.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)