StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
- Published
- Source
- arXiv
- Paper number
- 884
- Field
- Computer Vision
- arXiv ID
- 2608.12314
Key points
- It stores the world as an editable state composed of each object's 3D shape, position and pose, and semantic attributes.
- Front-view images are used for object appearance and top-view images for spatial layout, and a vision-language model resolves conflicts when the two disagree.
- Scene expansion, style changes, and object movement or pose changes are handled by editing parts of the state or swapping assets, which reduces the need to regenerate the whole scene.
- It proposes camera paths first and then checks rendered results to correct occlusion, bad composition, collision risk, and unstable motion.
- The average VBench score is 0.8484, the highest among the compared methods, and in a user study with 30 participants the combined scene and video scores were both 4.5 out of 5.
Paper links
External research summaries. These are not HDATF publications or measured product results.