From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models
TL;DR - VWG-Bench evaluates whether video generators can follow symbolic rules, physical constraints, and intentional goals rather than merely produce visually convincing footage. It exposes persistent reasoning failures and introduces Vid-PRE, a model-agnostic prompt rewriter that improves constraint adherence without modifying generators.
- VWG-Bench covers 9 reasoning dimensions and 38 fine-grained tasks.
- A three-level VLM-as-Judge protocol separately assesses visual fluency, rule adherence, and goal realization.
- Leading models render well but struggle with logic-heavy and rule-constrained tasks.
- Vid-PRE uses reinforcement learning with text-only rewards to produce concise, constraint-aware prompts and improves reasoning across multiple generators.