Scientific AI Videos Can Look Good and Still Be Wrong, Sci‑VBench Finds

Across 1,253 expert-annotated cases and 16 video models, perceptual scores were similar while scientific and causal correctness varied substantially.

Can the model animate scientific causality?

Sci‑VBench tests more than surface realism. Its prompts require scientific knowledge and causal reasoning over time. The authors assembled 1,253 expert-annotated cases spanning 60 subjects across natural science, healthcare, humanities and social science, and engineering.

They benchmarked 16 frontier proprietary and open-source video models. According to the paper, automatic perceptual-quality scores clustered relatively tightly, while prompt grounding and scientific and causal correctness varied substantially. The authors also report a pronounced proprietary–open gap. A smoother video is therefore not necessarily a more accurate simulation.

What educational producers should check

Scientific video needs correct forces, material changes, biological sequences and time scales—not just plausible shapes. Producers can turn source material into a list of verifiable events and ask a domain expert to review cause and effect across frames. Generated video should not serve as sole evidence in medical or safety instruction.

Research boundary

This is a preprint, not completed peer review, using 16 models and a new rubric. The authors report relatively high agreement between experts, non-experts and MLLM judges under their protocol, but that does not validate automated judging for every scientific domain.

Primary source