Beacon: Knowing When and How to Perform Agentic Visual Reasoning

Published
Source
arXiv
Paper number
771
Field
Computer Vision
arXiv ID
2607.28595

Key points

  • It defines two problems in tool use: Mode Adaptiveness, knowing when a tool is needed, and Tool Effect, knowing whether the tool actually helps.
  • Existing models gained little from tools because the gains on hard problems were offset by errors on easy ones.
  • The Necessity-Aware Adaptive Reward (NAAR) reduces unnecessary tool calls, and Hint-Conditioned Exploration (HCE) strengthens tool use on hard problems.
  • It achieves a best-in-class average of 58.98% across 13 benchmarks, a +6.07-point gain over the base model, and the largest tool gain-minus-gap improvement at +3.14 points.
  • It recovers about 40% of the all-wrong group through hint rollouts, effectively reusing the learning signal.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)