Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

Published
Source
arXiv
Paper number
1001
Field
Computer Vision
arXiv ID
2608.23329

Key points

  • It integrated video cropping (zooming in on a segment), image search, text search, and web browsing into a single exploration loop.
  • In situations such as news, demonstrations, documentaries, and lectures, where clues in a video must be connected to knowledge outside it, even small models can answer by filling knowledge gaps through retrieval.
  • It built an automated data pipeline that produces 26K verified SFT trajectories and 3K reinforcement-learning instances.
  • VideoRover-8B-RL scored 56.0% on VideoDR, comparable to GPT-5 without tools at 57.0%.
  • However, removing all external search tools caused VideoDR accuracy to collapse from 56.0% to 13.0%, showing that performance depends heavily on real-time web search.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)