Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement

Published
Source
arXiv
Paper number
674
Field
Computer Vision
arXiv ID
2607.17967

Key points

  • It diagnoses the limitations of existing monocular geometry estimation models as a structural mismatch rather than a capacity problem. When 3D information is decoded in a 2D coordinate system, features blend across depth boundaries.
  • The solution is to lift geometric modeling from the 2D image plane into 3D space, defining feature neighborhoods by 3D distance instead of screen distance.
  • It converts the point map into (X/Z, Y/Z, log Z) so that only log depth needs to be corrected, and quantizes this value to resolution D to build a sparse voxel shell.
  • A sparse 3D U-Net predicts log-depth residuals at each iteration, and the result redefines the voxel structure for the next iteration. Because of this self-guided scheme, computation is proportional to the number of occupied voxels.
  • In replacement experiments with a parameter-matched 2D U-Net, local metrics were similar to or worse than the base model, showing that the improvement comes from 3D inductive bias rather than capacity.
  • It was trained with K=3, but even when K was increased to 7 at inference time, the metrics held steady or improved slightly, indicating stable convergence of the learned correction.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)