Understanding Reasoning from Pretraining to Post-Training

Published
Source
arXiv
Paper number
657
Field
Machine Learning
arXiv ID
2607.16097

Key points

  • It runs 36 pretraining-RL combinations with models from 5M to 1B parameters and discovers a scaling law in which pretraining loss predicts post-RL performance.
  • RL increases the probability of existing correct answers on easy problems and surfaces answers that were almost absent during SFT on hard problems.
  • The result shows that allocating a larger share of the total compute budget to RL becomes optimal as the overall budget grows.
  • The same pattern appears not only in chess but also in math, across a 1B language model and 10B to 200B tokens, where longer pretraining improves RL efficiency.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)