MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- Published
- Source
- arXiv
- Paper number
- 749
- Field
- Computer Vision
- arXiv ID
- 2607.25948
Key points
- It is the first model to process 15 modalities with a single decoder, without modality-specific heads, losses, or task pipelines.
- Unlike prior any-to-any models, it uses a large pretrained decoder, BAGEL, to greatly reduce training cost.
- Uniform timestep sampling avoids modality confusion and enables stable training.
- It naturally supports chained generation, converting through intermediate modalities, and cross-modal self-verification, checking its own outputs through other modalities.
- It releases the 29-million-sample MODUS-DATASET and two checkpoints as open source.
Paper links
External research summaries. These are not HDATF publications or measured product results.