AdaCodec: A Predictive Visual Code for Video MLLMs
- Published
- Source
- arXiv
- Paper number
- 292
- Field
- Computer Vision
- arXiv ID
- 2606.02569
Key points
- However, existing video multimodal large language models, or video MLLMs, usually encode each sampled frame independently as an RGB image, which causes visual tokens to repeat content that already appears in earlier frames.
- Across 11 benchmarks, AdaCodec outperforms the frame-wise RGB baseline of Qwen3-VL-8B under the same visual-token budget.
- Even with one-seventh of the budget, AdaCodec with 32K tokens beats the 224K baseline on every long-form video benchmark, and on five general video benchmarks it raises the average score while reducing time to first token from 9.26 seconds to 1.62 seconds.
Paper links
External research summaries. These are not HDATF publications or measured product results.