3D-Aware VLMs with Implicit and Explicit Geometries

Published
Source
arXiv
Paper number
706
Field
Computer Vision
arXiv ID
2607.21595

Key points

  • We inject both implicit geometry tokens (IGT) and explicit geometry tokens (EGT) into the VLM at the same time.
  • EGT encodes detailed 3D information, such as depth maps and point clouds, into lightweight embeddings.
  • A 3D-aware adapter effectively fuses 2D vision, implicit 3D, and explicit 3D information.
  • It uses only RGB video as input, so no additional 3D sensors are needed.
  • It achieves SOTA on four tasks: 3D video detection, 3D visual grounding, 3D dense captioning, and spatial reasoning.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)