InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning
- Published
- Source
- arXiv
- Paper number
- 398
- Field
- Computer Vision
- arXiv ID
- 2606.12195
Key points
- MCR is a closed-loop multimodal reasoning formulation that integrates observation, reasoning, tool actions, feedback, and memory into a shared evolving context.
- M2LA, or Multimodal Multi-head Latent Attention, compresses the KV cache to make long multimodal rollouts efficient.
- The staged training pipeline includes M2LA conversion, continued pretraining, short-to-long video SFT, verifiable-task RL, and on-policy distillation.
- It shows substantial gains over InternVideo2.5-7B on long-video benchmarks such as Video-MME, MLVU, and EgoSchema.
- It instantiates a video agent with retrieval and verification tools to demonstrate the practicality of recursive multimodal reasoning.
- Because the open-weight ecosystem is evolving rapidly, the paper emphasizes the principled value of context efficiency and closed-loop reasoning over state-of-the-art claims.
Paper links
External research summaries. These are not HDATF publications or measured product results.