🛰️ Daily AI Frontier
‹ back to 2026-09-25

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

arXiv cs.CV Multimodal & Generative Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong 2026-09-24
Representative image for WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

TL;DR - WanPE is a 397B-parameter prompt-enhancement model that converts user requests into shot-level cinematic plans for text-to-video generation. Paired with Wan3.0, it substantially improves human preference, especially for 30-second videos.

  • Trained on 1.05M real-world videos using video-grounded reverse construction rather than forward prompt rewriting.
  • Semantic-Consistency GRPO preserves user requirements across shots and over time.
  • WanPEval covers 5–30-second videos and includes roughly 11K blind pairwise human assessments.
  • Human preference gains over raw prompts range from 10.66–18.84 points at 5–15 seconds to 50.86 points at 30 seconds.

view merged work →