Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Published
Source
arXiv
Paper number
830
Field
Computer Vision
arXiv ID
2608.02980

Key points

  • It compresses visual tokens spatially with 3D position information, including depth and camera pose, to process long videos efficiently.
  • With 3D Rotary Positional Embedding, attention operates directly in 3D space, which improves cross-view reasoning.
  • A structure where the language model and the segmentation decoder share the full visual token set greatly improves 3D visual grounding.
  • It outperforms existing 3D LMMs by 4% in 3D visual grounding at Acc@25 and by 13% in 3D instance segmentation mAP.
  • Its 2D vision-language performance stays at the level of Qwen2.5-VL, which shows that 3D training does not hurt 2D capability.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)