Linear Scaling Video VLMs for Long Video Understanding

Published
Source
arXiv
Paper number
284
Field
Computer Vision
arXiv ID
2605.31598

Key points

  • Existing efficiency methods improve scalability, but they often lose accuracy relative to full self-attention, for example through aggressive frame or token dropping or coarse attention approximations.
  • These results suggest practical progress toward scalable long-video understanding.
  • Video vision-language models, or VLMs, are increasingly used in long-horizon and streaming settings, but most video encoders still rely on spatiotemporal self-attention, so compute and latency grow quadratically with the number of frames.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)