Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Published
Source
arXiv
Paper number
041
Field
Reasoning / RL
arXiv ID
2503.01307

Key points

  • Some large language models achieve substantial self-improvement through reinforcement learning on verifiable reasoning tasks, while others quickly plateau under the same training.
  • It remained unclear which intrinsic properties or early capabilities enable LLMs to use extra compute effectively for self-improvement.
  • There was limited understanding of the foundational traits that determine success or failure in RL-based self-improvement for LLM reasoning.
  • The team developed a framework that uses a GPT-4o-mini classifier to identify and quantify four cognitive behaviors in LLM reasoning traces: verification, backtracking, subgoal setting, and reverse chaining.
  • They ran comparative RL experiments on models with different self-improvement abilities, Qwen-2.5-3B and Llama-3.2-3B, in the Countdown game.
  • They induced self-improvement with targeted interventions, such as fine-tuning models on synthetic data rich in specific cognitive behaviors and curating pretraining data to amplify those behaviors.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)