SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

Published
Source
arXiv
Paper number
372
Field
AI / General
arXiv ID
2606.09669

Key points

  • Design a simulator-agnostic evaluation protocol that integrates eight simulators.
  • Build a dataset of 760 tasks across six domains, including household chores, travel, social cooperation, and digital games.
  • Provide an MLLM-native interface with first-person vision-only observations and text-based actions.
  • Space reasoning remains a major challenge, with GPT-5 at 17.4% and Qwen-3.5 at 14.1%.
  • Domain-specific leaders differ: GPT-5 leads on household chores, while Gemini-3.1-Pro leads on games.
  • It finds a mismatch between task success rate and execution efficiency.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)