🛰️ Daily AI Frontier
‹ back to 2026-09-19

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

arXiv cs.LG Multimodal & Generative Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang, Michael Liu, Thomas Creavin, Kurt Keutzer, Xiuyu Li, Zhaoyang Lv, Chenfeng Xu, Haiwen Feng 2026-09-17
Representative image for Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

TL;DR - Video DeltaNet introduces a hybrid attention architecture for diffusion-based livestream video generation, combining local Softmax attention with bidirectional linear memory. On MiniMax H3, it accelerates 768p video denoising by 14.5Ă— while retaining fine-grained and long-range interactions.

  • Video Delta Attention updates its linear memory once per frame by jointly incorporating all spatial tokens.
  • Separate projections and learnable gates balance the local Softmax and long-range linear-attention branches.
  • A staged teacher-alignment recipe integrates the new pathway into pretrained models; text- and audio-related interactions retain Softmax attention.
  • With eight-step distillation and optimized SGLang serving, VDN-H3 denoises a 14.3-second video in 6.70 seconds on eight NVIDIA B200 GPUs.

view merged work →