Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- Published
- Source
- arXiv
- Paper number
- 752
- Field
- Computer Vision
- arXiv ID
- 2607.24904
Key points
- The codec-native tokenizer, Mage-ViT, reduces visual tokens by more than 75% by encoding only dynamic patches.
- Trained from scratch on just 560M unlabeled images, it matches SigLIP2, which uses billions of image-text pairs, demonstrating pretraining data efficiency.
- A dual-system design, with a System 1 event gate for fast reaction and a System 2 causal decoder for deep reasoning, supports real-time streaming understanding.
- With 4B parameters, it matches Qwen3-VL-4B on static tasks and surpasses even the 15B Phi-4 model on video understanding and 2D/3D spatial reasoning.
- It achieves up to 3.5x inference speedup and reports seven key empirical findings.
Paper links
External research summaries. These are not HDATF publications or measured product results.