Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning
- Published
- Source
- arXiv
- Paper number
- 402
- Field
- Computer Vision
- arXiv ID
- 2606.11683
Key points
- It reformulates spatial reasoning as a two-stage process of hypothesis and verification rather than a single turn.
- It strategically synthesizes complementary oblique views using VGGT monocular 3D reconstruction.
- It is a training-free framework that can be applied at inference time without modifying the MLLM architecture.
- It improves the VSI-Bench average by 5.2 points; at the sample level, the positive flip rate is 71.6% versus a negative flip rate of 28.4%, a 2.52 to 1 ratio.
- On an A100 GPU, it takes about 11 seconds per sample; reducing the VGGT input to 20 frames preserves a 2.8-point gain in about 4 seconds.
- Oblique Sweep outperforms bird's-eye views because it exposes spatial information while avoiding viewpoint-distribution mismatch.
Paper links
External research summaries. These are not HDATF publications or measured product results.