How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning

Published
Source
arXiv
Paper number
266
Field
Computer Vision
arXiv ID
2605.27310

Key points

  • The key failure mode of visual thinking is identified as "generated but not used."
  • View Dropout, or VDrop, is a training method that masks part of the input view during the answer segment to force the model to pass through a thinking image.
  • The Learnability-Informativeness, or L-I, tradeoff framework compares three visual-thinking representations: panorama, top-down, and point matching.
  • The model is trained on 8K synthetic Infinigen Indoors samples and evaluated on five real OOD benchmarks.
  • Only panorama visual thinking combined with VDrop scores highly on both axes and achieves the best OOD performance.
  • It achieves a 6.7-point OOD improvement over prior methods that were trained with three times more data.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)