Can the model animate scientific causality?
Sci‑VBench tests more than surface realism. Its prompts require scientific knowledge and causal reasoning over time. The authors assembled 1,253 expert-annotated cases spanning 60 subjects across natural science, healthcare, humanities and social science, and engineering.
They benchmarked 16 frontier proprietary and open-source video models. According to the paper, automatic perceptual-quality scores clustered relatively tightly, while prompt grounding and scientific and causal correctness varied substantially. The authors also report a pronounced proprietary–open gap. A smoother video is therefore not necessarily a more accurate simulation.
What educational producers should check
Scientific video needs correct forces, material changes, biological sequences and time scales—not just plausible shapes. Producers can turn source material into a list of verifiable events and ask a domain expert to review cause and effect across frames. Generated video should not serve as sole evidence in medical or safety instruction.
Research boundary
This is a preprint, not completed peer review, using 16 models and a new rubric. The authors report relatively high agreement between experts, non-experts and MLLM judges under their protocol, but that does not validate automated judging for every scientific domain.