Self-Play Pretraining with Zero Data
- Published
- Source
- arXiv
- Paper number
- 1117
- Field
- Machine Learning
- arXiv ID
- 2609.30063
Key points
- Proposed a pretraining paradigm where models generate their own training data through self-play instead of human curation
- The generator uses reinforcement learning to pick programs at the frontier of the learner's ability, forming an adaptive curriculum
- Without any natural data, zero-shot performance on natural language, images, speech, DNA, and more improved as a power law in compute
- The learner showed in-context learning on unseen tasks, and the generator discovered known mathematical sequences on its own
- An initial proof of extending the Bitter Lesson to training data itself: growing with compute alone, free of data curation bottlenecks
Paper links
External research summaries. These are not HDATF publications or measured product results.