rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
- Published
- Source
- arXiv
- Paper number
- 015
- Field
- Math Reasoning
- arXiv ID
- 2501.04519
Key points
- This method uses MCTS both at inference time and during data generation, with a math policy SLM proposing steps and an SLM-based Process Preference Model guiding the search.
- The core idea is to replace naive step-level reward annotations with preference pairs derived from MCTS Q-values, making the process reward model more robust than direct Q-value regression or outcome-only rewards.
- In four rounds of self-evolution, the system starts from a large DeepSeek-Coder bootstrap and progressively improves SLM-r1 through SLM-r4 and PPM-r1 through PPM-r4 using millions of verified solutions for 747K problems.
- With Qwen2.5-Math-7B, it reaches 90.0% on MATH and 53.3% on AIME, but the approach is compute-intensive and best suited to tasks where intermediate or final answers can be checked programmatically.
Paper links
External research summaries. These are not HDATF publications or measured product results.