World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible
- Published
- Source
- arXiv
- Paper number
- 407
- Field
- Computer Vision
- arXiv ID
- 2606.13652
Key points
- A pixel-aligned multilayer pointmap predicts L camera-space 3D points for each pixel, completing even occluded surfaces.
- WT-DiT combines a frozen MoGe encoder with three-way decomposed attention over layer, ray, and global dimensions.
- The depth-filling strategy forward-fills empty values in occluded layers and learns dense XYZ without mask prediction.
- It handles objects, scenes, and dynamic clips with a single architecture and does not require camera intrinsics as input.
- It outperforms existing depth predictors on visible surface accuracy and image-to-3D generators on Chamfer distance.
- It also supports text-based 3D scene editing, geometry-conditioned novel-view video synthesis, and training-free texture mesh generation.
Paper links
External research summaries. These are not HDATF publications or measured product results.