🛰️ Daily AI Frontier
‹ back to 2026-09-12

From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models

arXiv cs.CV Multimodal & Generative Meng Luo, Yicheng Liu, Jiahao Wang, Yuanxing Zhang, Xin Tao, Pengfei Wan, Kun Gai, Hao Fei 2026-09-10
Representative image for From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models

TL;DR - VWG-Bench evaluates whether video generators can follow symbolic rules, physical constraints, and intentional goals rather than merely produce visually convincing footage. It exposes persistent reasoning failures and introduces Vid-PRE, a model-agnostic prompt rewriter that improves constraint adherence without modifying generators.

  • VWG-Bench covers 9 reasoning dimensions and 38 fine-grained tasks.
  • A three-level VLM-as-Judge protocol separately assesses visual fluency, rule adherence, and goal realization.
  • Leading models render well but struggle with logic-heavy and rule-constrained tasks.
  • Vid-PRE uses reinforcement learning with text-only rewards to produce concise, constraint-aware prompts and improves reasoning across multiple generators.

view merged work →