PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

Published
Source
arXiv
Paper number
560
Field
Computer Vision
arXiv ID
2607.02515

Key points

  • It is the first pixel-space approach to apply diffusion directly to raw 3D point maps without going through a latent VAE space.
  • Direct x-prediction of the clean point map is far better than v-prediction, with Rel 35.44 versus 9.29.
  • Injecting DINOv3 4-layer features raises boundary sharpness to BF1 13.47, which is better than MoGe-2 at 11.75 and DA3 at 12.58.
  • The generative approach clearly beats deterministic regression on boundary quality, with BF1 13.92 versus 10.90, and on handling ambiguous regions.
  • It was accepted to ICML 2026 in Seoul, showing that pixel-space diffusion can extend beyond natural images to 3D and 4D geometry.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)