OmniReasoner: Thinking with Long Audio-Video via Native Tool Use
- Published
- Source
- arXiv
- Paper number
- 686
- Field
- Computer Vision
- arXiv ID
- 2607.19339
Key points
- We propose a tool-use learning framework that adaptively invokes zoom-in tools for long audio-video inputs.
- TimeAnchor preserves temporal coordinate consistency across different sampling resolutions.
- Temporal Augmented Data Engine automatically creates training data without manual effort.
- Accuracy improves from 29.3 to 34.8 on OmniVideoBench and from 32.0 to 35.4 on LVOmniBench.
- As videos get longer, the zoom-in tool is called more often, and we confirm that this contributes to improved accuracy.
Paper links
External research summaries. These are not HDATF publications or measured product results.