Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

Published
Source
arXiv
Paper number
813
Field
Computer Vision
arXiv ID
2608.03979

Key points

  • We designed a deep-research agent pipeline that selects key scenes from video frames and performs visual search.
  • We found and fixed modality bias, where the model prefers text search only, and knowledge leakage, where it answers from memory without tools.
  • After SFT, we used GRPO, Group Relative Policy Optimization, to build autonomous exploration ability beyond imitation learning.
  • We newly built VideoDR-Bench, with 200 complex multi-hop VQA problems.
  • The 35B-A3B model reaches SOTA at 64.0%, beating Claude-4.5-Sonnet at 59.0% by five points.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)