Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

Published
Source
arXiv
Paper number
780
Field
Machine Learning
arXiv ID
2607.27372

Key points

  • It discovers a third scaling axis in generative model training called exploration. In addition to parameters and data, more exploration improves performance.
  • It improves FLOP efficiency by 4.1x, sample efficiency by 6.2x, and parameter efficiency by 47%. The effect grows with data size, from 7% to 36%.
  • On ImageNet 256x256, it achieves an FID of 1.43 without guidance, near state of the art.
  • It enables end-to-end generative models, where inference matches training, and gets similar performance with 16x to 256x fewer steps than diffusion.
  • It observes that as exploration increases, generalization also improves.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)