SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

arXiv:2607.27084 · cs.CV, cs.AI · Submitted 2026-07-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context".

Jane: The paper was written by Zihan Deng, Chuanzhi Xu, Huiqi Liang, Haoyang Li, Xiaozhen Zhong et al. from The University of Hong Kong and The University of Sydney and University of Electronic Science and Technology of China.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: We are looking at a fascinating new paper called SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context.

Jane: That title is quite a mouthful, Tom, but it points to a massive problem in how we read science.

Tom: It really does, because usually, when people talk about image quality, they just want to know if a photo is blurry or bright.

Jane: Exactly, but this research from Zihan Deng and his team at the University of Hong Kong and the University of Sydney says that isn't enough for a scientific paper.

Tom: Right, because a graph could be crystal clear but still tell a complete lie about the data it's supposed to show.

Jane: That is why they added "Full-Manuscript Context" to the title, which means the AI has to look at the text around the image to see if they actually match.

Lu: This is such a creative leap because it moves AI from being a simple photographer to being a critical reader of evidence.

Tom: Do you think that's actually possible for a machine to do, Lu?

Lu: I think it's the only way we can scale up the verification of human knowledge as the volume of papers explodes.

Meng: I wonder how this would actually work in a real-world peer review pipeline, though.

Jane: Are you worried about the computational cost, Meng?

Meng: Not just the cost, but whether a researcher would actually trust an automated score to flag their work.

Lalam: Trust is the foundation of our entire scientific culture, and this tool could help protect that foundation.

Tom: That is a heavy thought, Lalam, but it's a vital one.

Jane: It really makes you realize that a picture isn't just worth a thousand words, it's actually tied to every single word in the paper.

Tom: We should look closer at what these researchers actually measured to see how they define "quality."

Summary: Tom: We've established that context is king, so let's look at how SciFigQual-Bench actually breaks down a scientific figure.

Jane: They don't just give one score, but instead use five different dimensions to judge an image.

Tom: You mean things like clarity and layout, right?

Jane: Yes, but they also include caption consistency, context consistency, and something called misleading risk.

Tom: Misleading risk sounds particularly important for preventing scientific fraud.

Jane: It is, because it checks for things like truncated axes or missing baselines that could trick a reader.

Meng: I was looking at their data scale, and they used seven thousand six hundred nine images from over a thousand papers.

Tom: That is a massive amount of data for a benchmark, isn't it?

Meng: It's huge, especially since they pulled these from top-tier computer science conferences between two thousand twenty and two thousand twenty-five.

Lu: The diversity is what excites me, since they cover everything from natural language processing to computer vision.

Jane: It means the benchmark isn't just biased toward one specific way of making charts.

Lu: It's like building a universal language for judging how visual evidence is presented.

Lalam: When we have a standard like this, it helps ensure that the visual language of science remains honest across all cultures.

Tom: It's a way to make sure the "visual truth" matches the "textual truth" everywhere.

Jane: We need to talk about the specific technology they built to handle all these moving parts.

Improvements: Tom: Now we are getting into the real engine of this paper, which is the SFQ-Agent.

Jane: This isn't just a single AI model looking at a picture, but a staged process that acts like a detective.

Tom: So it's not just a one-shot question to a model?

Jane: No, it collects visual evidence first, then language evidence, and finally fuses them together.

Meng: That staged approach is much smarter than just throwing a caption and an image at a standard model.

Tom: Why is that so much more effective, Meng?

Meng: Because it prevents the language model from just hallucinating a reason that sounds good but isn't actually in the pixels.

Lu: I love that it's "auditable," meaning you can actually trace the score back to the specific evidence found.

Jane: And the results they found were pretty incredible, especially with the GPT-five point six-Sol model.

Tom: They hit an average absolute error of only zero point four one eight, didn't they?

Jane: They did, and they also achieved a ninety-three point four percent consistency rate within one point of the human experts.

Meng: That level of accuracy is actually close to what a human reviewer would provide.

Lu: It opens up the possibility of having an AI assistant that helps researchers catch their own mistakes before they even submit a paper.

Lalam: This moves AI from being a generator of content to being a guardian of accuracy.

Tom: It really does, and it sets a new bar for what we should expect from multimodal models.

Jane: We've covered a lot of ground, so let's wrap this up.

Conclusion: Tom: This has been an intense look at SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context.

Jane: It's clear that the future of scientific AI isn't just about seeing, but about understanding the relationship between sight and text.

Tom: Before we head out, I want to hear one last thought from the team.

Lu: I see this as the first step toward an automated, global standard for scientific integrity.

Meng: From my side, I'm looking forward to seeing how these staged agents get integrated into actual software tools for engineers.

Lalam: This technology will ultimately help us build a more reliable digital archive of human achievement.

Tom: Thanks for joining us, everyone.

Jane: We'll see you next time for the next paper!

The University of Hong Kong · The University of Sydney · University of Electronic Science and Technology of China

cs.CV, cs.AI

Submitted: 2026-07-29

Updated: 2026-09-26

Comments: † Equal contribution. Affiliations: 1: The University of Hong Kong 2: The University of Sydney 3: University of Electronic Science and Technology of China Corresponding authors: Zihan Deng (zhdeng@hku.hk), Chuanzhi Xu (chuanzhi.xu@sydney.edu.au) Project page: https://frankdengai.github.io/SciFigQual-Bench Source code & dataset: https://github.com/FrankDengAI/SciFigQual-Bench

Code: https://github.com/FrankDengAI/SciFigQual-Bench

Project page: https://frankdengai.github.io/SciFigQual-Bench

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 80/100

The gist: The paper introduces SciFigQual-Bench, a rigorous benchmark designed to assess the quality of scientific figure interpretation by integrating full-manuscript context.

Key concepts

SciFigQual-Bench
This benchmark assesses scientific figures using five dimensions: clarity, layout, caption consistency, context consistency, and misleading risk. It was trained on 7,609 images from over a thousand papers to ensure that the assessment is not biased toward one specific type of chart.
Full-Manuscript Context
This methodology requires AI to look beyond just the image itself. By incorporating the surrounding text of the paper, it verifies if the visual evidence presented in a figure actually matches and supports the claims made in the written manuscript.
SFQ-Agent
This is a staged AI process that acts like a detective. Instead of one simple check, it collects visual evidence first, then language evidence, and finally fuses both to prevent models from hallucinating reasons that are not actually visible in the pixels.

Terminology

Summary

The paper introduces SciFigQual-Bench, a rigorous benchmark designed to assess the quality of scientific figure interpretation by integrating full-manuscript context. This advancement is critical because existing benchmarks often fail to capture the nuanced relationship between visual data presentation, accompanying textual descriptions (captions), and the overall narrative flow of a scientific manuscript. By establishing a comprehensive framework that evaluates multiple dimensions of understanding, SciFigQual-Bench aims to push the state-of-the-art in multimodal reasoning for scientific literature.

Benchmark Architecture and Scoring Dimensions

The evaluation system is built upon an extensive dataset comprising 1,200 instances from the eval1200 pool. Each record is evaluated across five distinct dimension scores, which allows for a granular understanding of failure modes. Furthermore, the system captures not only the numerical scores but also detailed per-rater reason strings and a summary/suggestion field. The scoring mechanism emphasizes accountability; Contested or low scores require reasons; unfaithful rationales are rejected in review. This methodology ensures that the assessment process itself is transparent and robust, ultimately leading to finalized equations (3)–(4) on the ACL 2025 holdout.

Human vs. Model Comparison Paradigms

The benchmark facilitates direct comparisons between human expert judgment and advanced AI models, including Direct GPT-5.6-Sol (D2), Direct Qwen-VL-Max (D4), and SFQ-Agent GPT-5.6-Sol (F3). The comparison is not merely holistic; it dissects specific failure points, such as:

  • Caption Bottleneck: As seen in Ex-G, models may over-credit fluent but incomplete captions, whereas the agentic approach (F3) is designed to check caption facts against image-side expectations.

  • Training-Curve Mismatch: In cases like Ex-H, the system can detect discrepancies where single-pass judges might soften contradictions, while advanced fusion techniques lower scores by combining the visual risk band with text severity (Algorithm 1).

Adjudication and Reliability Protocols

To maintain the highest level of scientific rigor, the benchmark incorporates a formal adjudication process. On the ACL 2025 holdout, any per-rater overall gaps above two points trigger adjudication before release. This systematic review ensures that discrepancies are resolved through expert consensus. The system utilizes multiple types of context for evaluation:

  1. Caption & Citing Text: Evaluating how well the figure is described in its immediate textual surroundings.

  2. Full Context (CTX): Assessing understanding based on the entire manuscript flow, as demonstrated by Ex-C where context explains inputs (X i), outputs (i), embeddings (P ij, R ij), and latent category vectors (h ij).

  3. Multi-Rater Consistency: The system reports dimension-first means, providing a comprehensive view of agreement across multiple human annotators.

Model Performance Analysis

The analysis highlights specific strengths and weaknesses across different model architectures. For instance, Ex-A demonstrates that while a caption broadly describes performance but omits Recall@K metrics and ADS/CDS groups visible in the legend, the detailed scoring mechanism can pinpoint these omissions. Furthermore, the comparison between models like D2, D4, and F3 on instances such as Ex-F shows strong tri-modal agreement, with advanced agents tracking gold scores within narrow margins (plus or minus 0.5 on all dimensions), while simpler models may slightly overscore the overall assessment.

Improvements for AI systems

This paper outlines sophisticated methodologies for evaluating complex, multimodal AI systems, particularly focusing on grounding, coherence, and cross-modal consistency. The core strength lies not just in what is evaluated (Captioning/Citing), but how the evaluation is structured (multi-rater consensus, dimension gating, agentic refinement).

Based on this rigorous framework, I can propose improvements across three major areas: Evaluation Pipeline Architecture, Model Grounding & Reasoning, and Agentic Refinement.


The Improvement: Current models often rely on single, monolithic scoring mechanisms (e.g., one overall coherence score). The system must be upgraded to process and output scores for every defined dimension independently before calculating an aggregated score. This requires a specialized Meta-Evaluator Module.

Technical Specifications:

  • Dimension Separation: Force the model to explicitly separate scoring into distinct, weighted dimensions (e.g., Visual Fidelity (VC), Semantic Linkage (SL), Caption Completeness (CC), Contextual Accuracy (CTX), Misleading Risk (MR)).

  • Gating Mechanism: Implement a gated null mechanism (as seen in Ex-A/Ex-C) where the system must identify and explicitly report when a dimension is not applicable or cannot be inferred from the provided context, preventing arbitrary zero-scoring.

  • Adjudication Trigger: Integrate an internal consistency checker that flags any scoring discrepancy between dimensions (e.g., if VC is high but SL is low) and forces a self-correction or a detailed rationale generation before final output.

Improved AI Capability: The system moves from merely providing an answer to providing a fully traceable, multi-faceted critique. It can diagnose why its own output failed (e.g., The caption is semantically accurate (SL=9/10) but fails on visual fidelity because it omits the secondary data points (VC=6/10)—a level of diagnostic depth currently lacking).

Abstract

Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers. The few existing studies on scholarly charts remain confined to visual-surface comparisons, failing to verify caption alignment, citation relevance, or visual misleadingness. To address this, we propose SciFigQual-Bench, a full-text contextual benchmark that evaluates scientific images across five dimensions (clarity, layout, caption fit, context relevance, and misleading risk). The data covers top computer-science conferences from 2020 to 2025; 6,308 images were independently scored by multiple domain experts in five dimensions and aggregated into gold-standard annotations. Unlike previous scientific figure benchmarks, our dataset binds each image to its caption, citing sentence, and manuscript context. To enable automated evaluation on this benchmark, we designed a staged cross-modal evaluation framework SFQ-Agent to achieve auditable and refined scoring through the collection and fusion of modal evidence. Multiple mainstream large models were evaluated on the test subset eval1200, and SFQ-Agent (F3) equipped with GPT-5.6-Sol achieved the lowest overall average absolute error (0.418) and the highest consistency rate (93.4%), consistently outperforming both direct evaluation and auxiliary (Sidecar) visual language model evaluation schemes.

Sources

Related papers