SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
- Published
- Source
- arXiv
- Paper number
- 413
- Field
- Computer Vision
- arXiv ID
- 2606.13673
Key points
- The agent loop is designed around a stateful Python kernel that lets the system observe intermediate results and modify code at each step.
- It overcomes the limitations of both single-pass code with pre-commitment and structured tool calls with restricted composition.
- Across 20 spatial reasoning benchmarks, including static and dynamic 3D and 4D settings, it reaches 59.9 percent on average, which is 11.2 points higher than prior spatial agents.
- It improves performance consistently across six VLM backbones, from Qwen 27B to 397B and Gemma4, without benchmark-specific or model-specific tuning.
- The largest gains appear in dynamic 4D video reasoning and multi-view inference.
- It is training-free and is achieved only through the action-interface design, without changing the backbone model or system prompt.
Paper links
External research summaries. These are not HDATF publications or measured product results.