Scaling Native Multimodal Pre-Training From Scratch
- Published
- Source
- arXiv
- Paper number
- 721
- Field
- LLMs / NLP
- arXiv ID
- 2607.22043
Key points
- The scaling law for language objectives is almost unaffected by data composition, but multimodal objectives are very sensitive to the data ratio.
- Text-heavy data is more efficient for larger models, so the optimal allocation of resources shifts toward model capacity.
- It derives, in formula form, an efficiency frontier that specifies the optimal combination of model size, token count, and data ratio.
- It confirms a cross-modal transfer effect in which native multimodal training also improves pure text-space reasoning ability.
- It observes the natural emergence of multimodal in-context learning in a 3B-parameter model.
Paper links
External research summaries. These are not HDATF publications or measured product results.