Scalable Visual Pretraining for Language Intelligence

Published
Source
arXiv
Paper number
595
Field
Computer Vision
arXiv ID
2607.09657

Key points

  • The paper proposes a VP framework that pretrains directly from raw visual patches, without parsing documents into text, through next-latent prediction.
  • It shows consistent gains over TP across four backbones, with up to +3.22 on GPQA Diamond and +2.1 on MMLU-Pro.
  • It reaches equal or better performance with only 25 percent of the token budget, which shows that visual representations are more compression-efficient than text.
  • The gains are largest on visually dense documents such as figures, equations, and tables, which indicates recovery of information that text conversion would lose.
  • VP is effective not only for multimodal models but also for language-only models, which shows that it is architecture-agnostic.
  • A frozen vision encoder plus a trainable projection layer are enough to integrate the visual stream into an autoregressive backbone.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)