🛰️ Daily AI Frontier
‹ back to 2026-09-19

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Research Multimodal & Generative

Ranking

Overall 86
Content 95
Popularity 64

Observed public metrics from 1 member.

Representative image for Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Merged summary

TL;DR - Video DeltaNet introduces a hybrid attention architecture for diffusion-based livestream video generation, combining local Softmax attention with bidirectional linear memory. On MiniMax H3, it accelerates 768p video denoising by 14.5Ă— while retaining fine-grained and long-range interactions.

  • Video Delta Attention updates its linear memory once per frame by jointly incorporating all spatial tokens.
  • Separate projections and learnable gates balance the local Softmax and long-range linear-attention branches.
  • A staged teacher-alignment recipe integrates the new pathway into pretrained models; text- and audio-related interactions retain Softmax attention.
  • With eight-step distillation and optimized SGLang serving, VDN-H3 denoises a 14.3-second video in 6.70 seconds on eight NVIDIA B200 GPUs.

Sources (1)

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

arXiv cs.LG Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang, Michael Liu, Thomas Creavin, Kurt Keutzer, Xiuyu Li, Zhaoyang Lv, Chenfeng Xu, Haiwen Feng 2026-09-17 arXiv:2609.20744
Public signals Hugging Face upvotes 50
Providers: Hugging Face · Upvotes 50 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:18:05.118711 UTC

TL;DR - Video DeltaNet introduces a hybrid attention architecture for diffusion-based livestream video generation, combining local Softmax attention with bidirectional linear memory. On MiniMax H3, it accelerates 768p video denoising by 14.5Ă— while retaining fine-grained and long-range interactions.

  • Video Delta Attention updates its linear memory once per frame by jointly incorporating all spatial tokens.
  • Separate projections and learnable gates balance the local Softmax and long-range linear-attention branches.
  • A staged teacher-alignment recipe integrates the new pathway into pretrained models; text- and audio-related interactions retain Softmax attention.
  • With eight-step distillation and optimized SGLang serving, VDN-H3 denoises a 14.3-second video in 6.70 seconds on eight NVIDIA B200 GPUs.
item →