Evidence, not just visual polish
SciFigQual-Bench, submitted July 29, does not judge scientific figures like ordinary photographs. It links 6,308 images from major 2020–2025 computer-science conferences to their captions, citing sentences, and manuscript context, with independent ratings from multiple domain experts.
Five dimensions and initial results
The dimensions are clarity, layout, caption fit, contextual relevance, and misleading risk. The authors also propose SFQ-Agent, which gathers and combines multimodal evidence in stages.
On a 1,200-item test subset, the authors report that SFQ-Agent F3 with GPT-5.6 Sol achieved a mean absolute error of 0.418 and 93.4% consistency, the best among the compared large-model methods. They attribute the gain to reading the figure with its caption and manuscript instead of judging pixels alone.
Use and limitations
Researchers can use such systems as a preflight check for missing labels, caption-text mismatches, baselines, and misleading presentation—not as the final reviewer. People must still confirm that source data and code support the visual conclusion.
The data centers on computer-science conferences, and the authors evaluated their own staged system on the new benchmark. It is a preprint; generalization to other disciplines, languages, and publication formats remains unproven.