Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs
- Published
- Source
- arXiv
- Paper number
- 890
- Field
- Computer Vision
- arXiv ID
- 2608.12179
Key points
- It processes RGB frames from a short time window in order and applies the scale ratio predicted by the 3D reconstruction model to create object boxes in real-world units.
- It removes the step that converts 2D detections into 3D and directly predicts boxes in 3D space by combining geometric information from multiple viewpoints.
- On CA-1M, the AP15 of the single-frame baseline was 11.7, and using time information, camera information, and learnable multi-view processing raised it to 21.2.
- Without any extra training, it achieved AP15 of 15.2 on ScanNet200, surpassing the previous monocular best of 11.7.
- On CA-1M, it outperformed existing monocular and multi-view image methods, but it remained below methods that use real depth or 3D point data. It is also limited to detection that does not distinguish between indoor scenes and object categories.
Paper links
External research summaries. These are not HDATF publications or measured product results.