EchoWM: Open and Enterable Omnimodal World Models
- Published
- Source
- arXiv
- Paper number
- 993
- Field
- Computer Vision
- arXiv ID
- 2608.23189
Key points
- It is an omnimodal world model in which a single model simultaneously generates video, environmental sounds, background music, and speech.
- It unifies discrete commands and continuous poses into a common 6-DoF trajectory, handling first-person and third-person views without separate controllers.
- It resolved the conflict between quality and control through 4 stages: AV pretraining, control training, joint fine-tuning, and autoregressive post-training.
- It ranked first on the public WBench Navigation benchmark and also achieved top-tier quality on SANA-WM-Bench.
- However, without persistent three-dimensional memory, scene structure, characters, and sounds can gradually drift across repeated sequential generations.
Paper links
External research summaries. These are not HDATF publications or measured product results.