🛰️ Daily AI Frontier
‹ back to 2026-08-11

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

arXiv cs.CV Multimodal & Generative Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao 2026-08-10
Representative image for Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

TL;DR - Sci-VBench is a 1,253-example expert-annotated benchmark testing whether video generation models can produce scientifically correct, reasoning-grounded videos across 60 subjects in four disciplines. It matters because it shows visual realism gains have not translated into reliable modeling of scientific and causal dynamics.

  • Covers Natural Science, Healthcare, Humanities & Social Sciences, and Engineering, requiring temporally rich videos that need knowledge-grounded synthesis rather than surface-level plausibility.
  • Introduces a rubric-based evaluation protocol; both non-expert human raters and MLLM-as-Judge systems reach relatively high agreement with expert judgments, enabling reproducible scaled evaluation.
  • Across 16 frontier proprietary and open-source models, automatic perceptual-quality scores cluster tightly, while Prompt Grounding and Scientific/Causal Correctness vary widely.
  • A pronounced proprietary–open-source gap emerges specifically on the reasoning-sensitive dimensions.

view merged work →