Wonder: Video World Model Done Better
- Published
- Source
- arXiv
- Paper number
- 745
- Field
- Computer Vision
- arXiv ID
- 2607.26037
Key points
- It converts camera motion into rendered visual cues rather than abstract data, which prevents control accuracy from collapsing after distillation.
- A sparse full-fidelity memory preserves every past frame while keeping attention cost constant.
- A mixture-of-students over time and camera-aware adversarial regularization preserve both visual quality and control accuracy during distillation.
- It achieves meaningful improvements over the strongest baseline on rotation position error, going from 0.0174 to 0.0132 for translation and from 0.1155 to 0.0784 for rotation.
- It also supports video input, so it can re-shoot existing videos by changing the camera path in real time.
Paper links
External research summaries. These are not HDATF publications or measured product results.