Finetuning with Sampling: SFT Learns Better Than You Think
- Published
- Source
- arXiv
- Paper number
- 1148
- Field
- Machine Learning
- arXiv ID
- 2610.02140
Key points
- They proposed projection sampling, which rewrites expert answers via MCMC sampling to be closer to the model's own expressions before SFT.
- On chemistry tasks, plain SFT lost substantial MMLU-level capability, while this method forgot the least with an average loss of only -1.10%.
- Increasing MCMC steps moved the training data closer to the model distribution while accuracy rose in tandem, showing sampling itself is a scalable compute axis.
- The trained model acquired genuinely new abilities, solving problems the base model could never solve with up to 67.2% pass rate.
- Initializing RL on top of sampling SFT achieved the best overall results with +40.7% on MATH500 and +25.1% on GSM8K.
Paper links
External research summaries. These are not HDATF publications or measured product results.