Phi-4-reasoning-vision-15B Technical Report
- Published
- Source
- arXiv
- Paper number
- 128
- Field
- Multimodal / Reasoning
- arXiv ID
- 2603.03975
Key points
- The dominant trend in vision-language models (VLMs) is to keep increasing model size, which brings huge training and inference costs and hinders broad deployment.
- Many existing multimodal models struggle with tasks that require fine-grained visual detail, such as understanding user interfaces for computer-use agents, because their perceptual ability is limited.
- There is a need for practical insights and open-weight models that can achieve efficient and capable multimodal reasoning without demanding enormous resources.
- We developed Phi-4-reasoning-vision-15B, a 15-billion-parameter open-weight multimodal model with an intermediate-fusion architecture that uses the SigLIP-2 vision encoder and the Phi-4-Reasoning LLM backbone.
- We implemented a three-stage training recipe consisting of MLP pretraining, extensive instruction tuning on diverse and carefully curated datasets, and a final stage for long context, multiple images, and responsible AI (RAI) data.
- We emphasized data quality through systematic filtering, AI-assisted error correction, and synthetic augmentation, and integrated a dynamic-resolution vision encoder to optimize high-resolution visual perception.
Paper links
External research summaries. These are not HDATF publications or measured product results.