S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence
- Published
- Source
- arXiv
- Paper number
- 454
- Field
- Computer Vision
- arXiv ID
- 2606.20515
Key points
- It reframes spatial reasoning as the accumulation of spatiotemporal evidence rather than frame-by-frame prediction.
- It uses a three-layer tool stack: 2D grounding, 3D geometric lifting, and spatial knowledge aggregation.
- Scene Memory, which stores object state, and Agent Memory, which stores reasoning history, support stateful reasoning.
- In a training-free setting, it improves MMSI-Bench by 4.5 percent over GPT-5.4.
- S-AGENT-8B trained on 300K trajectories improves by 10.5 percentage points over Qwen3-VL-8B, from 31.1 percent to 41.6 percent.
- The key idea is that a level-3 expert transforms raw 3D data into task-oriented measurements.
Paper links
External research summaries. These are not HDATF publications or measured product results.