WorldSculpt: Generating Compositional Worlds from Grounded Videos
- Published
- Source
- arXiv
- Paper number
- 1073
- Field
- Computer Vision
- arXiv ID
- 2609.05416
Key points
- By adding a circuit that feeds multi-view video into a single-object 3D generative model (Pixal3D), it handles scenes with hundreds of objects without scene-level training.
- Each object is produced as an individual mesh and placed in world coordinates, so games, AR, simulation, and robotics can pick objects out separately.
- It releases UE-MeshyScene, an Unreal Engine benchmark of 6 scenes with 93 to 701 objects and ground-truth meshes per object.
- On the heavily occluded benchmark, learning-based image-based rendering cuts CD-L2 error by 12% and raises F-Score from 0.944 to 0.951.
- It also demonstrates, without extra training, converting 3DGS scenes from generative world models such as Marble and HY-World 2.0 into object-level mesh scenes.
Paper links
External research summaries. These are not HDATF publications or measured product results.