Cosmos 3: Omnimodal World Models for Physical AI
- Published
- Source
- arXiv
- Paper number
- 347
- Field
- Computer Vision
- arXiv ID
- 2606.02800
Key points
- It is an omnimodal design based on a single mixture-of-transformers architecture that flexibly supports input and output across language, image, video, audio, and action.
- A dual-tower layer structure and dual-stream joint attention effectively fuse multimodal information.
- It serves both as a reasoner and a generator, and is released in Super and Nano scales.
- It uses large synthetic datasets, such as SDG-PhyxSim, SDG-RobotSim, and SDG-DriveSim, for training.
- According to Artificial Analysis, it was rated the best open-source text-to-image and image-to-video model, and the top policy model in RoboArena.
- It fully releases code, checkpoints, datasets, and evaluation benchmarks under the Linux Foundation OpenMDW-1.1 license.
Paper links
External research summaries. These are not HDATF publications or measured product results.