Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

Published
Source
arXiv
Paper number
821
Field
Computer Vision
arXiv ID
2608.05000

Key points

  • It finds a strong asymmetry in which language and visual understanding act as priors that drive visual generation, while the reverse transfer is weak.
  • When data complexity is low, modalities are synergistic, but when complexity is high they compete, and a shared attention structure promotes synergy.
  • Early joint training is far more effective than late alignment, and late fusion causes a vision-laziness effect.
  • Using these principles, the authors create a training recipe that uses only 5 percent of the compute to train a 13.5B MoE model on 2T tokens and still achieves strong generation quality.
  • The synergy pattern remains consistent regardless of visual tokenizer design, which shows that architecture choice is the main driver.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)