Scalable Visual Pretraining for Language Intelligence
- Published
- Source
- arXiv
- Paper number
- 595
- Field
- Computer Vision
- arXiv ID
- 2607.09657
Key points
- The paper proposes a VP framework that pretrains directly from raw visual patches, without parsing documents into text, through next-latent prediction.
- It shows consistent gains over TP across four backbones, with up to +3.22 on GPQA Diamond and +2.1 on MMLU-Pro.
- It reaches equal or better performance with only 25 percent of the token budget, which shows that visual representations are more compression-efficient than text.
- The gains are largest on visually dense documents such as figures, equations, and tables, which indicates recovery of information that text conversion would lose.
- VP is effective not only for multimodal models but also for language-only models, which shows that it is architecture-agnostic.
- A frozen vision encoder plus a trainable projection layer are enough to integrate the visual stream into an autoregressive backbone.
Paper links
External research summaries. These are not HDATF publications or measured product results.