Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- Published
- Source
- arXiv
- Paper number
- 780
- Field
- Machine Learning
- arXiv ID
- 2607.27372
Key points
- It discovers a third scaling axis in generative model training called exploration. In addition to parameters and data, more exploration improves performance.
- It improves FLOP efficiency by 4.1x, sample efficiency by 6.2x, and parameter efficiency by 47%. The effect grows with data size, from 7% to 36%.
- On ImageNet 256x256, it achieves an FID of 1.43 without guidance, near state of the art.
- It enables end-to-end generative models, where inference matches training, and gets similar performance with 16x to 256x fewer steps than diffusion.
- It observes that as exploration increases, generalization also improves.
Paper links
External research summaries. These are not HDATF publications or measured product results.