Beacon: Knowing When and How to Perform Agentic Visual Reasoning
- Published
- Source
- arXiv
- Paper number
- 771
- Field
- Computer Vision
- arXiv ID
- 2607.28595
Key points
- It defines two problems in tool use: Mode Adaptiveness, knowing when a tool is needed, and Tool Effect, knowing whether the tool actually helps.
- Existing models gained little from tools because the gains on hard problems were offset by errors on easy ones.
- The Necessity-Aware Adaptive Reward (NAAR) reduces unnecessary tool calls, and Hint-Conditioned Exploration (HCE) strengthens tool use on hard problems.
- It achieves a best-in-class average of 58.98% across 13 benchmarks, a +6.07-point gain over the base model, and the largest tool gain-minus-gap improvement at +3.14 points.
- It recovers about 40% of the all-wrong group through hint rollouts, effectively reusing the learning signal.
Paper links
External research summaries. These are not HDATF publications or measured product results.