Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
- Published
- Source
- arXiv
- Paper number
- 830
- Field
- Computer Vision
- arXiv ID
- 2608.02980
Key points
- It compresses visual tokens spatially with 3D position information, including depth and camera pose, to process long videos efficiently.
- With 3D Rotary Positional Embedding, attention operates directly in 3D space, which improves cross-view reasoning.
- A structure where the language model and the segmentation decoder share the full visual token set greatly improves 3D visual grounding.
- It outperforms existing 3D LMMs by 4% in 3D visual grounding at Acc@25 and by 13% in 3D instance segmentation mAP.
- Its 2D vision-language performance stays at the level of Qwen2.5-VL, which shows that 3D training does not hurt 2D capability.
Paper links
External research summaries. These are not HDATF publications or measured product results.