Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

Published
Source
arXiv
Paper number
775
Field
Computer Vision
arXiv ID
2607.28611

Key points

  • It designs a hybrid diffusion backbone that combines KDA, linear attention, MLA, global attention, and unit consistency for local patterns.
  • Without positional embeddings, it processes text, images, and video in a single stream using only raster order.
  • HeteroP automatically transfers hyperparameters across scales based on functional fan-in for each tensor.
  • With 11B total parameters and 2B active parameters, it achieves 7.3x better compute efficiency than Wan-2.1 at 2B, and training takes about 600 H100 days.
  • It is trained on 5-second clips and can extrapolate to 30-second video generation in zero shot, with only a 6 percent FID drop at 6x longer length.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)