MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

Published
Source
arXiv
Paper number
749
Field
Computer Vision
arXiv ID
2607.25948

Key points

  • It is the first model to process 15 modalities with a single decoder, without modality-specific heads, losses, or task pipelines.
  • Unlike prior any-to-any models, it uses a large pretrained decoder, BAGEL, to greatly reduce training cost.
  • Uniform timestep sampling avoids modality confusion and enables stable training.
  • It naturally supports chained generation, converting through intermediate modalities, and cross-modal self-verification, checking its own outputs through other modalities.
  • It releases the 29-million-sample MODUS-DATASET and two checkpoints as open source.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)