🛰️ Daily AI Frontier
‹ back to 2026-09-12

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

arXiv cs.CV Multimodal & Generative Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Deyuan Liu, Jungang Li, Dechuang Chen, Ming Lin, Jingjiang Zhou, Haopeng Jin, Qi Jia, Xiaohang Wang, Yaole Wang, Zhanqiang Zhang, Ran Li, Zhengkun Huang, Shuyue Xiong, Yuji Wang, Zikun Dai, Hui He, Yang Luo, Mang Ning, Weiqi Feng, Chengyang Ye, Xinyue Lin, Min Zhao, Hongzhou Zhu, Hengkai Tan, Zeyuan Wang, Chendong Xiang, Kaiwen Zheng, Zhijie Deng, Fan Bao, Jianfei Chen, Jun Zhu 2026-09-10
Representative image for Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

TL;DR - Vidu S2 introduces models for real-time 720p digital-character generation and live video editing, with an exploration of spatial video generation. It advances interactive generative video by allowing references and visual elements to be changed during a stream.

  • Vidu S2-Avatar supports dynamic references, stronger instruction following, and real-time 720p generation.
  • Vidu S2-Editing enables live style rendering and replacement of clothing, characters, or backgrounds.
  • The authors explore real-time spatial video generation for both avatar generation and video editing.
  • Reported experiments show Vidu S2 outperforming the evaluated baselines.

view merged work →