WorldSculpt: Generating Compositional Worlds from Grounded Videos

Published
Source
arXiv
Paper number
1073
Field
Computer Vision
arXiv ID
2609.05416

Key points

  • By adding a circuit that feeds multi-view video into a single-object 3D generative model (Pixal3D), it handles scenes with hundreds of objects without scene-level training.
  • Each object is produced as an individual mesh and placed in world coordinates, so games, AR, simulation, and robotics can pick objects out separately.
  • It releases UE-MeshyScene, an Unreal Engine benchmark of 6 scenes with 93 to 701 objects and ground-truth meshes per object.
  • On the heavily occluded benchmark, learning-based image-based rendering cuts CD-L2 error by 12% and raises F-Score from 0.944 to 0.951.
  • It also demonstrates, without extra training, converting 3DGS scenes from generative world models such as Marble and HY-World 2.0 into object-level mesh scenes.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)