rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Published
Source
arXiv
Paper number
015
Field
Math Reasoning
arXiv ID
2501.04519

Key points

  • This method uses MCTS both at inference time and during data generation, with a math policy SLM proposing steps and an SLM-based Process Preference Model guiding the search.
  • The core idea is to replace naive step-level reward annotations with preference pairs derived from MCTS Q-values, making the process reward model more robust than direct Q-value regression or outcome-only rewards.
  • In four rounds of self-evolution, the system starts from a large DeepSeek-Coder bootstrap and progressively improves SLM-r1 through SLM-r4 and PPM-r1 through PPM-r4 using millions of verified solutions for 747K problems.
  • With Qwen2.5-Math-7B, it reaches 90.0% on MATH and 53.3% on AIME, but the approach is compute-intensive and best suited to tasks where intermediate or final answers can be checked programmatically.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)