Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

Published
Source
arXiv
Paper number
873
Field
Computer Vision
arXiv ID
2608.09873

Key points

  • We built 1,253 expert-annotated video-generation tasks across 60 courses in four fields: natural sciences, medicine, humanities and social sciences, and engineering.
  • We provide an expert-written evaluation rubric and step-by-step scoring criteria from 1 to 5 for every example.
  • Non-expert raters also achieve high agreement with experts when given the rubric, with Cohen's kappa = 0.842.
  • Evaluation of 16 frontier models reveals similar visual quality but large gaps in scientific accuracy.
  • Prompt improvements help somewhat, but the limits in spatial and temporal consistency come from the generator itself.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)