Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
- Published
- Source
- arXiv
- Paper number
- 821
- Field
- Computer Vision
- arXiv ID
- 2608.05000
Key points
- It finds a strong asymmetry in which language and visual understanding act as priors that drive visual generation, while the reverse transfer is weak.
- When data complexity is low, modalities are synergistic, but when complexity is high they compete, and a shared attention structure promotes synergy.
- Early joint training is far more effective than late alignment, and late fusion causes a vision-laziness effect.
- Using these principles, the authors create a training recipe that uses only 5 percent of the compute to train a 13.5B MoE model on 2T tokens and still achieves strong generation quality.
- The synergy pattern remains consistent regardless of visual tokenizer design, which shows that architecture choice is the main driver.
Paper links
External research summaries. These are not HDATF publications or measured product results.