Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 079
- Field
- Reasoning / RL
- arXiv ID
- 2509.03646
Key points
- Despite empirical success, the mechanistic understanding of how reinforcement learning (RL) improves complex reasoning in LLMs remained unclear.
- Puzzling phenomena during RL training, such as the aha moment and length-scaling effects, lacked a unified explanation.
- Existing RL algorithms often apply optimization pressure indiscriminately to every generated token, which leads to inefficient learning.
- We functionally decomposed LLM-generated tokens into high-level planning tokens, such as logical maneuvers, and low-level execution tokens, such as calculations.
- We introduced Strategic Grams (SGs), a data-driven method that empirically identifies planning tokens without subjective manual annotation.
- We developed Hierarchy-Aware Credit Assignment (HICRA), an RL algorithm that concentrates optimization by amplifying advantage signals on identified planning tokens.
Paper links
External research summaries. These are not HDATF publications or measured product results.