OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

Published
Source
arXiv
Paper number
686
Field
Computer Vision
arXiv ID
2607.19339

Key points

  • We propose a tool-use learning framework that adaptively invokes zoom-in tools for long audio-video inputs.
  • TimeAnchor preserves temporal coordinate consistency across different sampling resolutions.
  • Temporal Augmented Data Engine automatically creates training data without manual effort.
  • Accuracy improves from 29.3 to 34.8 on OmniVideoBench and from 32.0 to 35.4 on LVOmniBench.
  • As videos get longer, the zoom-in tool is called more often, and we confirm that this contributes to improved accuracy.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)