3D-Aware VLMs with Implicit and Explicit Geometries
- Published
- Source
- arXiv
- Paper number
- 706
- Field
- Computer Vision
- arXiv ID
- 2607.21595
Key points
- We inject both implicit geometry tokens (IGT) and explicit geometry tokens (EGT) into the VLM at the same time.
- EGT encodes detailed 3D information, such as depth maps and point clouds, into lightweight embeddings.
- A 3D-aware adapter effectively fuses 2D vision, implicit 3D, and explicit 3D information.
- It uses only RGB video as input, so no additional 3D sensors are needed.
- It achieves SOTA on four tasks: 3D video detection, 3D visual grounding, 3D dense captioning, and spatial reasoning.
Paper links
External research summaries. These are not HDATF publications or measured product results.