SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
- Published
- Source
- arXiv
- Paper number
- 372
- Field
- AI / General
- arXiv ID
- 2606.09669
Key points
- Design a simulator-agnostic evaluation protocol that integrates eight simulators.
- Build a dataset of 760 tasks across six domains, including household chores, travel, social cooperation, and digital games.
- Provide an MLLM-native interface with first-person vision-only observations and text-based actions.
- Space reasoning remains a major challenge, with GPT-5 at 17.4% and Qwen-3.5 at 14.1%.
- Domain-specific leaders differ: GPT-5 leads on household chores, while Gemini-3.1-Pro leads on games.
- It finds a mismatch between task success rate and execution efficiency.
Paper links
External research summaries. These are not HDATF publications or measured product results.