AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research

Published
Source
arXiv
Paper number
1026
Field
Computer Vision
arXiv ID
2608.25559

Key points

  • Addressing the fact that the optimal tool combination differs by question and video type, it lets the model skip tool calls based on its own capabilities and internal knowledge.
  • It proposed 'model-conditional tool necessity filtering,' which removes calls from training trajectories when the target model can answer without tools.
  • It built and released the VDR-EE benchmark and training data, covering 250 entity- and event-centered questions across 7 domains.
  • It achieved the best performance among the open-source models evaluated and also showed substantial improvement over the base model on VideoDR.
  • Redundancy-aware reward RL reduced average tool calls from 7.84 to 6.80 while maintaining accuracy (51%).

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)