SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

Published
Source
arXiv
Paper number
699
Field
Computer Vision
arXiv ID
2607.21553

Key points

  • It mixes linear attention and softmax attention in a 3-to-1 ratio so that it can achieve both speed and quality.
  • The AttnRes technique improves information quality in deep layers by about 12 percent.
  • A single H100 GPU can generate a 720p video in 13 seconds, which is 120 times faster than Wan 2.2.
  • It directly learns the optimal balance point by showing that 25 percent softmax attention is the best trade-off between quality and efficiency.
  • Kernel optimization through Sol-Engine delivers an additional 3.58 times speedup.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)