Finetuning with Sampling: SFT Learns Better Than You Think

Published
Source
arXiv
Paper number
1148
Field
Machine Learning
arXiv ID
2610.02140

Key points

  • They proposed projection sampling, which rewrites expert answers via MCMC sampling to be closer to the model's own expressions before SFT.
  • On chemistry tasks, plain SFT lost substantial MMLU-level capability, while this method forgot the least with an average loss of only -1.10%.
  • Increasing MCMC steps moved the training data closer to the model distribution while accuracy rose in tandem, showing sampling itself is a scalable compute axis.
  • The trained model acquired genuinely new abilities, solving problems the base model could never solve with up to 67.2% pass rate.
  • Initializing RL on top of sampling SFT achieved the best overall results with +40.7% on MATH500 and +25.1% on GSM8K.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)