🛰️ Daily AI Frontier
‹ back to 2026-07-24

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

arXiv cs.CV Multimodal & Generative Junsong Chen, Jincheng Yu, Yitong Li, Shuchen Xue, Haozhe Liu, Jingyu Xin, Yuyang Zhao, Tian Ye, Zhangjie Wu, Zian Wang, Daquan Zhou, Ping Luo, Song Han, Enze Xie 2026-07-23

TL;DR - SANA-Video 2.0 is a 5B/14B hybrid-attention video diffusion transformer that targets softmax-level quality with linear-attention efficiency. It enables high-quality 720p generation on a single GPU while substantially reducing latency.

  • Combines gated linear attention with periodic softmax anchors at a 3:1 ratio to preserve efficient long-sequence scaling and full-rank interactions.
  • Block Attention Residuals reuse anchor features across layers, increasing deep-layer effective rank by about 12%.
  • The 14B model scores 84.30 on VBench in 13.2 seconds at 480p using 40 sampling steps on one H100.
  • A compiled forward pass is 3.2× faster than matched full-softmax attention at 720p/60s; additional Sol-Engine optimizations accelerate the backbone by 3.58×.

view merged work →