AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research
- Published
- Source
- arXiv
- Paper number
- 1026
- Field
- Computer Vision
- arXiv ID
- 2608.25559
Key points
- Addressing the fact that the optimal tool combination differs by question and video type, it lets the model skip tool calls based on its own capabilities and internal knowledge.
- It proposed 'model-conditional tool necessity filtering,' which removes calls from training trajectories when the target model can answer without tools.
- It built and released the VDR-EE benchmark and training data, covering 250 entity- and event-centered questions across 7 domains.
- It achieved the best performance among the open-source models evaluated and also showed substantial improvement over the base model on VideoDR.
- Redundancy-aware reward RL reduced average tool calls from 7.84 to 6.80 while maintaining accuracy (51%).
Paper links
External research summaries. These are not HDATF publications or measured product results.