Self-Play Pretraining with Zero Data

Published
Source
arXiv
Paper number
1117
Field
Machine Learning
arXiv ID
2609.30063

Key points

  • Proposed a pretraining paradigm where models generate their own training data through self-play instead of human curation
  • The generator uses reinforcement learning to pick programs at the frontier of the learner's ability, forming an adaptive curriculum
  • Without any natural data, zero-shot performance on natural language, images, speech, DNA, and more improved as a power law in compute
  • The learner showed in-context learning on unseen tasks, and the generator discovered known mathematical sequences on its own
  • An initial proof of extending the Bitter Lesson to training data itself: growing with compute alone, free of data curation bottlenecks

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)