EchoWM: Open and Enterable Omnimodal World Models

Published
Source
arXiv
Paper number
993
Field
Computer Vision
arXiv ID
2608.23189

Key points

  • It is an omnimodal world model in which a single model simultaneously generates video, environmental sounds, background music, and speech.
  • It unifies discrete commands and continuous poses into a common 6-DoF trajectory, handling first-person and third-person views without separate controllers.
  • It resolved the conflict between quality and control through 4 stages: AV pretraining, control training, joint fine-tuning, and autoregressive post-training.
  • It ranked first on the public WBench Navigation benchmark and also achieved top-tier quality on SANA-WM-Bench.
  • However, without persistent three-dimensional memory, scene structure, characters, and sounds can gradually drift across repeated sequential generations.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)