Wonder: Video World Model Done Better

Published
Source
arXiv
Paper number
745
Field
Computer Vision
arXiv ID
2607.26037

Key points

  • It converts camera motion into rendered visual cues rather than abstract data, which prevents control accuracy from collapsing after distillation.
  • A sparse full-fidelity memory preserves every past frame while keeping attention cost constant.
  • A mixture-of-students over time and camera-aware adversarial regularization preserve both visual quality and control accuracy during distillation.
  • It achieves meaningful improvements over the strongest baseline on rotation position error, going from 0.0174 to 0.0132 for translation and from 0.1155 to 0.0784 for rotation.
  • It also supports video input, so it can re-shoot existing videos by changing the camera path in real time.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)