Understanding Reasoning from Pretraining to Post-Training
- Published
- Source
- arXiv
- Paper number
- 657
- Field
- Machine Learning
- arXiv ID
- 2607.16097
Key points
- It runs 36 pretraining-RL combinations with models from 5M to 1B parameters and discovers a scaling law in which pretraining loss predicts post-RL performance.
- RL increases the probability of existing correct answers on easy problems and surfaces answers that were almost absent during SFT on hard problems.
- The result shows that allocating a larger share of the total compute budget to RL becomes optimal as the overall budget grows.
- The same pattern appears not only in chess but also in math, across a 1B language model and 10B to 200B tokens, where longer pretraining improves RL efficiency.
Paper links
External research summaries. These are not HDATF publications or measured product results.