Phi-4-reasoning-vision-15B Technical Report

Published
Source
arXiv
Paper number
128
Field
Multimodal / Reasoning
arXiv ID
2603.03975

Key points

  • The dominant trend in vision-language models (VLMs) is to keep increasing model size, which brings huge training and inference costs and hinders broad deployment.
  • Many existing multimodal models struggle with tasks that require fine-grained visual detail, such as understanding user interfaces for computer-use agents, because their perceptual ability is limited.
  • There is a need for practical insights and open-weight models that can achieve efficient and capable multimodal reasoning without demanding enormous resources.
  • We developed Phi-4-reasoning-vision-15B, a 15-billion-parameter open-weight multimodal model with an intermediate-fusion architecture that uses the SigLIP-2 vision encoder and the Phi-4-Reasoning LLM backbone.
  • We implemented a three-stage training recipe consisting of MLP pretraining, extensive instruction tuning on diverse and carefully curated datasets, and a final stage for long context, multiple images, and responsible AI (RAI) data.
  • We emphasized data quality through systematic filtering, AI-assisted error correction, and synthetic augmentation, and integrated a dynamic-resolution vision encoder to optimize high-resolution visual perception.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)