Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation

Published
Source
arXiv
Paper number
422
Field
Computer Vision
arXiv ID
2606.17030

Key points

  • It standardizes more than 20 embodiments and more than 500 action categories into a single language interface by using natural language as a unified action interface.
  • A double-stream MMDiT with 60 layers fuses semantic information from Qwen2.5-VL and video-VAE latents through layer-wise joint attention.
  • The EWK dataset contains 8.6 million video-text pairs and more than 200 million frames, including 5.9 million manipulation episodes, 200K autonomous-driving episodes, and more than 6K navigation episodes.
  • It ranks first on EWMBench with 4.60, ahead by 33 percent over HSD, first on DreamGen Bench with 4.952, first among open-source models on WorldModelBench with 8.99, and scores 0.804 on PBench.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)