Linear Scaling Video VLMs for Long Video Understanding
- Published
- Source
- arXiv
- Paper number
- 284
- Field
- Computer Vision
- arXiv ID
- 2605.31598
Key points
- Existing efficiency methods improve scalability, but they often lose accuracy relative to full self-attention, for example through aggressive frame or token dropping or coarse attention approximations.
- These results suggest practical progress toward scalable long-video understanding.
- Video vision-language models, or VLMs, are increasingly used in long-horizon and streaming settings, but most video encoders still rely on spatiotemporal self-attention, so compute and latency grow quadratically with the number of frames.
Paper links
External research summaries. These are not HDATF publications or measured product results.