Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement
- Published
- Source
- arXiv
- Paper number
- 674
- Field
- Computer Vision
- arXiv ID
- 2607.17967
Key points
- It diagnoses the limitations of existing monocular geometry estimation models as a structural mismatch rather than a capacity problem. When 3D information is decoded in a 2D coordinate system, features blend across depth boundaries.
- The solution is to lift geometric modeling from the 2D image plane into 3D space, defining feature neighborhoods by 3D distance instead of screen distance.
- It converts the point map into (X/Z, Y/Z, log Z) so that only log depth needs to be corrected, and quantizes this value to resolution D to build a sparse voxel shell.
- A sparse 3D U-Net predicts log-depth residuals at each iteration, and the result redefines the voxel structure for the next iteration. Because of this self-guided scheme, computation is proportional to the number of occupied voxels.
- In replacement experiments with a parameter-matched 2D U-Net, local metrics were similar to or worse than the base model, showing that the improvement comes from 3D inductive bias rather than capacity.
- It was trained with K=3, but even when K was increased to 7 at inference time, the metrics held steady or improved slightly, indicating stable convergence of the learned correction.
Paper links
External research summaries. These are not HDATF publications or measured product results.