S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

Published
Source
arXiv
Paper number
454
Field
Computer Vision
arXiv ID
2606.20515

Key points

  • It reframes spatial reasoning as the accumulation of spatiotemporal evidence rather than frame-by-frame prediction.
  • It uses a three-layer tool stack: 2D grounding, 3D geometric lifting, and spatial knowledge aggregation.
  • Scene Memory, which stores object state, and Agent Memory, which stores reasoning history, support stateful reasoning.
  • In a training-free setting, it improves MMSI-Bench by 4.5 percent over GPT-5.4.
  • S-AGENT-8B trained on 300K trajectories improves by 10.5 percentage points over Qwen3-VL-8B, from 31.1 percent to 41.6 percent.
  • The key idea is that a level-3 expert transforms raw 3D data into task-oriented measurements.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)