Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
TL;DR - Sci-VBench is a 1,253-example expert-annotated benchmark testing whether video generation models can produce scientifically correct, reasoning-grounded videos across 60 subjects in four disciplines. It matters because it shows visual realism gains have not translated into reliable modeling of scientific and causal dynamics.
- Covers Natural Science, Healthcare, Humanities & Social Sciences, and Engineering, requiring temporally rich videos that need knowledge-grounded synthesis rather than surface-level plausibility.
- Introduces a rubric-based evaluation protocol; both non-expert human raters and MLLM-as-Judge systems reach relatively high agreement with expert judgments, enabling reproducible scaled evaluation.
- Across 16 frontier proprietary and open-source models, automatic perceptual-quality scores cluster tightly, while Prompt Grounding and Scientific/Causal Correctness vary widely.
- A pronounced proprietary–open-source gap emerges specifically on the reasoning-sensitive dimensions.