How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
- Published
- Source
- arXiv
- Paper number
- 266
- Field
- Computer Vision
- arXiv ID
- 2605.27310
Key points
- The key failure mode of visual thinking is identified as "generated but not used."
- View Dropout, or VDrop, is a training method that masks part of the input view during the answer segment to force the model to pass through a thinking image.
- The Learnability-Informativeness, or L-I, tradeoff framework compares three visual-thinking representations: panorama, top-down, and point matching.
- The model is trained on 8K synthetic Infinigen Indoors samples and evaluated on five real OOD benchmarks.
- Only panorama visual thinking combined with VDrop scores highly on both axes and achieves the best OOD performance.
- It achieves a 6.7-point OOD improvement over prior methods that were trained with three times more data.
Paper links
External research summaries. These are not HDATF publications or measured product results.