Cosmos 3: Omnimodal World Models for Physical AI

Published
Source
arXiv
Paper number
347
Field
Computer Vision
arXiv ID
2606.02800

Key points

  • It is an omnimodal design based on a single mixture-of-transformers architecture that flexibly supports input and output across language, image, video, audio, and action.
  • A dual-tower layer structure and dual-stream joint attention effectively fuse multimodal information.
  • It serves both as a reasoner and a generator, and is released in Super and Nano scales.
  • It uses large synthetic datasets, such as SDG-PhyxSim, SDG-RobotSim, and SDG-DriveSim, for training.
  • According to Artificial Analysis, it was rated the best open-source text-to-image and image-to-video model, and the top policy model in RoboArena.
  • It fully releases code, checkpoints, datasets, and evaluation benchmarks under the Linux Foundation OpenMDW-1.1 license.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)