World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible

Published
Source
arXiv
Paper number
407
Field
Computer Vision
arXiv ID
2606.13652

Key points

  • A pixel-aligned multilayer pointmap predicts L camera-space 3D points for each pixel, completing even occluded surfaces.
  • WT-DiT combines a frozen MoGe encoder with three-way decomposed attention over layer, ray, and global dimensions.
  • The depth-filling strategy forward-fills empty values in occluded layers and learns dense XYZ without mask prediction.
  • It handles objects, scenes, and dynamic clips with a single architecture and does not require camera intrinsics as input.
  • It outperforms existing depth predictors on visible surface accuracy and image-to-3D generators on Chamfer distance.
  • It also supports text-based 3D scene editing, geometry-conditioned novel-view video synthesis, and training-free texture mesh generation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)