AdaCodec: A Predictive Visual Code for Video MLLMs

Published
Source
arXiv
Paper number
292
Field
Computer Vision
arXiv ID
2606.02569

Key points

  • However, existing video multimodal large language models, or video MLLMs, usually encode each sampled frame independently as an RGB image, which causes visual tokens to repeat content that already appears in earlier frames.
  • Across 11 benchmarks, AdaCodec outperforms the frame-wise RGB baseline of Qwen3-VL-8B under the same visual-token budget.
  • Even with one-seventh of the budget, AdaCodec with 32K tokens beats the 224K baseline on every long-form video benchmark, and on five general video benchmarks it raises the average score while reducing time to first token from 9.26 seconds to 1.62 seconds.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)