From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - VWG-Bench evaluates whether video generators can follow symbolic rules, physical constraints, and intentional goals rather than merely produce visually convincing footage. It exposes persistent reasoning failures and introduces Vid-PRE, a model-agnostic prompt rewriter that improves constraint adherence without modifying generators.
- VWG-Bench covers 9 reasoning dimensions and 38 fine-grained tasks.
- A three-level VLM-as-Judge protocol separately assesses visual fluency, rule adherence, and goal realization.
- Leading models render well but struggle with logic-heavy and rule-constrained tasks.
- Vid-PRE uses reinforcement learning with text-only rewards to produce concise, constraint-aware prompts and improves reasoning across multiple generators.
Sources (1)
From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - VWG-Bench evaluates whether video generators can follow symbolic rules, physical constraints, and intentional goals rather than merely produce visually convincing footage. It exposes persistent reasoning failures and introduces Vid-PRE, a model-agnostic prompt rewriter that improves constraint adherence without modifying generators.
- VWG-Bench covers 9 reasoning dimensions and 38 fine-grained tasks.
- A three-level VLM-as-Judge protocol separately assesses visual fluency, rule adherence, and goal realization.
- Leading models render well but struggle with logic-heavy and rule-constrained tasks.
- Vid-PRE uses reinforcement learning with text-only rewards to produce concise, constraint-aware prompts and improves reasoning across multiple generators.