Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
- Published
- Source
- arXiv
- Paper number
- 873
- Field
- Computer Vision
- arXiv ID
- 2608.09873
Key points
- We built 1,253 expert-annotated video-generation tasks across 60 courses in four fields: natural sciences, medicine, humanities and social sciences, and engineering.
- We provide an expert-written evaluation rubric and step-by-step scoring criteria from 1 to 5 for every example.
- Non-expert raters also achieve high agreement with experts when given the rubric, with Cohen's kappa = 0.842.
- Evaluation of 16 frontier models reveals similar visual quality but large gaps in scientific accuracy.
- Prompt improvements help somewhat, but the limits in spatial and temporal consistency come from the generator itself.
Paper links
External research summaries. These are not HDATF publications or measured product results.