VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Published
Source
arXiv
Paper number
640
Field
Computer Vision
arXiv ID
2607.14935

Key points

  • The 3D vision Transformer, I3D-ViT, processes frames as spatiotemporal bundles rather than separate images, which greatly reduces token count.
  • An adaptive mechanism that adjusts frame resolution by importance improves efficiency for real-time streaming.
  • It releases 3 million high-quality multimodal instruction examples, including Academic2M, LV116K, and OL617K.
  • With a 4B model, it outperforms previous open-source models of the same class on 12 benchmarks, including MotionBench, VideoMME, and LVBench.
  • It releases the model, code, training strategy, dataset, and pipeline, which provides a reproducible foundation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)