🛰️ Daily AI Frontier
‹ back to 2026-09-09

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

Research Multimodal & Generative

Ranking

Overall 80
Content 85
Popularity 67

Observed public metrics from 1 member.

Merged summary

TL;DR - Mask Forcing improves autoregressive video diffusion distillation by perturbing student rollouts with spatial-temporal masks that mix cleaner signals into noisy inputs. This aims to reduce mode collapse, error accumulation, over-saturation, and over-smoothing without real video data or extra post-training.

  • Targets reverse-KL mode-seeking in Distribution Matching Distillation, which can collapse the student onto a limited subset of teacher modes.
  • Applies random spatial and temporal masks during self-rollout to encourage broader exploration of the teacher distribution.
  • Uses cleaner tokens to guide denoising of noisier tokens and improve intermediate rollout predictions.
  • Reportedly improves visual quality across multiple autoregressive video diffusion distillation methods while remaining efficient.

Sources (1)

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

arXiv cs.CV Zhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong, Songchun Zhang, Junchao Huang, Guian Fang, Xin Wang, Pan Hui, Anyi Rao 2026-09-08 arXiv:2609.09123
Public signals Hugging Face upvotes 54
Providers: Hugging Face · Upvotes 54 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:22:09.604435 UTC

TL;DR - Mask Forcing improves autoregressive video diffusion distillation by perturbing student rollouts with spatial-temporal masks that mix cleaner signals into noisy inputs. This aims to reduce mode collapse, error accumulation, over-saturation, and over-smoothing without real video data or extra post-training.

  • Targets reverse-KL mode-seeking in Distribution Matching Distillation, which can collapse the student onto a limited subset of teacher modes.
  • Applies random spatial and temporal masks during self-rollout to encourage broader exploration of the teacher distribution.
  • Uses cleaner tokens to guide denoising of noisier tokens and improve intermediate rollout predictions.
  • Reportedly improves visual quality across multiple autoregressive video diffusion distillation methods while remaining efficient.
item →