Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
- Published
- Source
- arXiv
- Paper number
- 813
- Field
- Computer Vision
- arXiv ID
- 2608.03979
Key points
- We designed a deep-research agent pipeline that selects key scenes from video frames and performs visual search.
- We found and fixed modality bias, where the model prefers text search only, and knowledge leakage, where it answers from memory without tools.
- After SFT, we used GRPO, Group Relative Policy Optimization, to build autonomous exploration ability beyond imitation learning.
- We newly built VideoDR-Bench, with 200 complex multi-hop VQA problems.
- The 35B-A3B model reaches SOTA at 64.0%, beating Claude-4.5-Sonnet at 59.0% by five points.
Paper links
External research summaries. These are not HDATF publications or measured product results.