SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

Published
Source
arXiv
Paper number
413
Field
Computer Vision
arXiv ID
2606.13673

Key points

  • The agent loop is designed around a stateful Python kernel that lets the system observe intermediate results and modify code at each step.
  • It overcomes the limitations of both single-pass code with pre-commitment and structured tool calls with restricted composition.
  • Across 20 spatial reasoning benchmarks, including static and dynamic 3D and 4D settings, it reaches 59.9 percent on average, which is 11.2 points higher than prior spatial agents.
  • It improves performance consistently across six VLM backbones, from Qwen 27B to 397B and Gemma4, without benchmark-specific or model-specific tuning.
  • The largest gains appear in dynamic 4D video reasoning and multi-view inference.
  • It is training-free and is achieved only through the action-interface design, without changing the backbone model or system prompt.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)