Scaling Native Multimodal Pre-Training From Scratch

Published
Source
arXiv
Paper number
721
Field
LLMs / NLP
arXiv ID
2607.22043

Key points

  • The scaling law for language objectives is almost unaffected by data composition, but multimodal objectives are very sensitive to the data ratio.
  • Text-heavy data is more efficient for larger models, so the optimal allocation of resources shifts toward model capacity.
  • It derives, in formula form, an efficiency frontier that specifies the optimal combination of model size, token count, and data ratio.
  • It confirms a cross-modal transfer effect in which native multimodal training also improves pure text-space reasoning ability.
  • It observes the natural emergence of multimodal in-context learning in a 3B-parameter model.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)