SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
- Published
- Source
- arXiv
- Paper number
- 913
- Field
- Computer Vision
- arXiv ID
- 2608.14138
Key points
- Existing systems often handle 3D reconstruction and spatial question answering with different models or separate modules.
- SPARGen tokenizes camera poses and text answers, while generating depth, point maps, and optical flow as images.
- It trains on spatial reasoning, visual geometry, and optical flow data, and it is trained for 100k steps on 64 H100 GPUs.
- In public model comparisons, it ranks first on 13 of 15 spatial reasoning subcategories. In the KITTI optical flow evaluation without additional fine-tuning, it records an EPE of 4.09 and an F1-all of 13.34.
- A specialized 3D model, VGGT, performs better on some reconstruction metrics, but it predicts normalized depth and point maps, so it cannot recover true metric scale in meters.
Paper links
External research summaries. These are not HDATF publications or measured product results.