WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
TL;DR - WanPE is a 397B-parameter prompt-enhancement model that converts user requests into shot-level cinematic plans for text-to-video generation. Paired with Wan3.0, it substantially improves human preference, especially for 30-second videos.
- Trained on 1.05M real-world videos using video-grounded reverse construction rather than forward prompt rewriting.
- Semantic-Consistency GRPO preserves user requirements across shots and over time.
- WanPEval covers 5–30-second videos and includes roughly 11K blind pairwise human assessments.
- Human preference gains over raw prompts range from 10.66–18.84 points at 5–15 seconds to 50.86 points at 30 seconds.