Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- Published
- Source
- arXiv
- Paper number
- 041
- Field
- Reasoning / RL
- arXiv ID
- 2503.01307
Key points
- Some large language models achieve substantial self-improvement through reinforcement learning on verifiable reasoning tasks, while others quickly plateau under the same training.
- It remained unclear which intrinsic properties or early capabilities enable LLMs to use extra compute effectively for self-improvement.
- There was limited understanding of the foundational traits that determine success or failure in RL-based self-improvement for LLM reasoning.
- The team developed a framework that uses a GPT-4o-mini classifier to identify and quantify four cognitive behaviors in LLM reasoning traces: verification, backtracking, subgoal setting, and reverse chaining.
- They ran comparative RL experiments on models with different self-improvement abilities, Qwen-2.5-3B and Llama-3.2-3B, in the Countdown game.
- They induced self-improvement with targeted interventions, such as fine-tuning models on synthetic data rich in specific cognitive behaviors and curating pretraining data to amplify those behaviors.
Paper links
External research summaries. These are not HDATF publications or measured product results.