Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
- Published
- Source
- arXiv
- Paper number
- 422
- Field
- Computer Vision
- arXiv ID
- 2606.17030
Key points
- It standardizes more than 20 embodiments and more than 500 action categories into a single language interface by using natural language as a unified action interface.
- A double-stream MMDiT with 60 layers fuses semantic information from Qwen2.5-VL and video-VAE latents through layer-wise joint attention.
- The EWK dataset contains 8.6 million video-text pairs and more than 200 million frames, including 5.9 million manipulation episodes, 200K autonomous-driving episodes, and more than 6K navigation episodes.
- It ranks first on EWMBench with 4.60, ahead by 33 percent over HSD, first on DreamGen Bench with 4.952, first among open-source models on WorldModelBench with 8.99, and scores 0.804 on PBench.
Paper links
External research summaries. These are not HDATF publications or measured product results.