Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Published
Source
arXiv
Paper number
752
Field
Computer Vision
arXiv ID
2607.24904

Key points

  • The codec-native tokenizer, Mage-ViT, reduces visual tokens by more than 75% by encoding only dynamic patches.
  • Trained from scratch on just 560M unlabeled images, it matches SigLIP2, which uses billions of image-text pairs, demonstrating pretraining data efficiency.
  • A dual-system design, with a System 1 event gate for fast reaction and a System 2 causal decoder for deep reasoning, supports real-time streaming understanding.
  • With 4B parameters, it matches Qwen3-VL-4B on static tasks and surpasses even the 15B Phi-4 model on video understanding and 2D/3D spatial reasoning.
  • It achieves up to 3.5x inference speedup and reports seven key empirical findings.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)