Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
- Published
- Source
- arXiv
- Paper number
- 775
- Field
- Computer Vision
- arXiv ID
- 2607.28611
Key points
- It designs a hybrid diffusion backbone that combines KDA, linear attention, MLA, global attention, and unit consistency for local patterns.
- Without positional embeddings, it processes text, images, and video in a single stream using only raster order.
- HeteroP automatically transfers hyperparameters across scales based on functional fan-in for each tensor.
- With 11B total parameters and 2B active parameters, it achieves 7.3x better compute efficiency than Wan-2.1 at 2B, and training takes about 600 H100 days.
- It is trained on 5-second clips and can extrapolate to 30-second video generation in zero shot, with only a 6 percent FID drop at 6x longer length.
Paper links
External research summaries. These are not HDATF publications or measured product results.