Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents
- Published
- Source
- arXiv
- Paper number
- 1001
- Field
- Computer Vision
- arXiv ID
- 2608.23329
Key points
- It integrated video cropping (zooming in on a segment), image search, text search, and web browsing into a single exploration loop.
- In situations such as news, demonstrations, documentaries, and lectures, where clues in a video must be connected to knowledge outside it, even small models can answer by filling knowledge gaps through retrieval.
- It built an automated data pipeline that produces 26K verified SFT trajectories and 3K reinforcement-learning instances.
- VideoRover-8B-RL scored 56.0% on VideoDR, comparable to GPT-5 without tools at 57.0%.
- However, removing all external search tools caused VideoDR accuracy to collapse from 56.0% to 13.0%, showing that performance depends heavily on real-time web search.
Paper links
External research summaries. These are not HDATF publications or measured product results.