Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
- Published
- Source
- arXiv
- Paper number
- 839
- Field
- LLMs / NLP
- arXiv ID
- 2608.05139
Key points
- It defines skill entropy as a directional metric that measures how hard it is to switch between two skills.
- It builds Skill2-Bench, a benchmark with 558 skills across 9 domains.
- Frontier models show monotonically decreasing accuracy on tasks with high skill entropy.
- Skill entropy RL improves Qwen3-4B from 34.4% to 68.4% and Qwen3-1.7B from 14.6% to 40.1%.
- It is a reusable training signal that can also be applied to existing datasets such as OpenR1-Math.
Paper links
External research summaries. These are not HDATF publications or measured product results.