VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
- Published
- Source
- arXiv
- Paper number
- 640
- Field
- Computer Vision
- arXiv ID
- 2607.14935
Key points
- The 3D vision Transformer, I3D-ViT, processes frames as spatiotemporal bundles rather than separate images, which greatly reduces token count.
- An adaptive mechanism that adjusts frame resolution by importance improves efficiency for real-time streaming.
- It releases 3 million high-quality multimodal instruction examples, including Academic2M, LV116K, and OL617K.
- With a 4B model, it outperforms previous open-source models of the same class on 12 benchmarks, including MotionBench, VideoMME, and LVBench.
- It releases the model, code, training strategy, dataset, and pipeline, which provides a reproducible foundation.
Paper links
External research summaries. These are not HDATF publications or measured product results.