Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

Published
Source
arXiv
Paper number
839
Field
LLMs / NLP
arXiv ID
2608.05139

Key points

  • It defines skill entropy as a directional metric that measures how hard it is to switch between two skills.
  • It builds Skill2-Bench, a benchmark with 558 skills across 9 domains.
  • Frontier models show monotonically decreasing accuracy on tasks with high skill entropy.
  • Skill entropy RL improves Qwen3-4B from 34.4% to 68.4% and Qwen3-1.7B from 14.6% to 40.1%.
  • It is a reusable training signal that can also be applied to existing datasets such as OpenR1-Math.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)