SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing

arXiv:2602.09809 · cs.CV, cs.AI · Submitted 2026-02-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing".

Jane: Scientific diagrams convey explicit structural information, yet modern text-to-image models often produce visually plausible but structurally incorrect results.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Let's talk about the title and who put this paper together. The researchers are Tong Zhang, Honglin Lin, Zhou Liu, Chong Chen, Wentao Zhang—a solid team from Peking University and Shanghai Jiao Tong University with some strong affiliations in cloud and data intelligence.

Jane: It’s a focused set of authors for a deep dive into scientific diagram generation because they clearly have expertise across the whole pipeline they describe.

Lu: The title, "SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing," tells us immediately that the core idea is using inverse parsing to test structure awareness in scientific diagrams.

Meng: So, instead of just asking if a diagram looks right, they are building a system to check if the underlying logic can be recovered from the pixel output.

Lalam: That makes sense because it moves us away from subjective image similarity and toward an objective measure of how well the AI understands scientific relationships.

The paper's summary: Tom: Moving on, they summarize SciFlow-Bench as a structure-first benchmark that evaluates models directly from pixel-level outputs by turning those images back into structured graphs to measure structural recoverability.

Jane: That’s the core idea, Tom: it’s not just about generating a picture; it’s about whether that picture holds the correct scientific logic when parsed.

Lu: The paper points out this "visual illusion" where diagrams can look perfect but have reversed dependencies or missing steps, and SciFlow-Bench is designed specifically to catch those errors.

Meng: It sounds like they are setting up a rigorous testing environment that forces the AI to prove it understands the functional components and their directed relations.

Lalam: I see this as a major step because we can start training models with explicit structural supervision derived from canonical ground-truth graphs, ensuring they adhere to predefined scientific workflows.

The paper's improvements: Tom: Now for the improvements they propose in "SciFlow-Bench," which involves a whole closed-loop, round-trip protocol powered by a hierarchical multiagent system that coordinates planning, perception, and structural reasoning.

Jane: That multiagent setup is pretty clever because it breaks down the complex task of diagram generation into specialized steps, from understanding text to finally generating the symbolic representation.

Lu: The Cognitive Planning layer converting method descriptions into a "structured visual prompt" seems like a vital bridge between natural language and the visual model itself, which is where things often get messy.

Meng: I’m looking at that Fine-Grained Perception Layer, with agents like the Environment Curator and Shape Hunter working together to generate grounded nodes before they even get to the reasoning stage.

Lalam: The Topology Coder emitting Mermaid code as a symbolic intermediate representation is something I think is very powerful because it allows us to enforce explicit connectivity decisions during the generation process.

Conclusion: Tom: So, wrapping things up with the conclusion of "SciFlow-Bench," they establish structural recoverability as the central criterion for diagram quality, suggesting that visual plausibility alone doesn't guarantee structural correctness.

Jane: That’s a strong statement; it means we can finally diagnose those invisible logical failures that image-centric metrics completely miss.

Lu: It really solidifies the need for a structure-first evaluation axis if we want to build multimodal systems that actually reason about scientific structure rather than just mimic the look of diagrams.

Meng: For practical engineering, this means we can develop structure-aware loss functions that heavily weight graph-level accuracy, especially edge-level F1 scores, during training.

Lalam: I think this entire framework is a principled and scalable way to diagnose structural failures that are otherwise invisible to standard visual metrics when we're looking at scientific diagrams.

Peking University · Shanghai Jiao Tong University

cs.CV, cs.AI

Submitted: 2026-02-10

Updated: 2026-06-09

Code: https://github.com/Tong-0302/SciFlow-Bench

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 87/100

The gist: Scientific diagrams convey explicit structural information, yet modern text-to-image models often produce visually plausible but structurally incorrect results.

Key concepts

Structure-First Criterion
This is an evaluation method where success depends entirely on whether the generated diagram preserves the logical relationships between its parts, rather than how visually appealing it looks. It focuses on whether connections are correctly made and dependencies are accurate, ignoring superficial visual qualities.
Closed-Loop, Round-Trip Protocol
This is the testing cycle used to evaluate models. First, a model creates an image from text; second, that image is analyzed (inverse-parsed) to create a predicted graph. This predicted graph is then compared directly against the known correct graph to measure structural recovery.
Hierarchical Multi-Agent System (HMAS)
This is the complex system used to perform the evaluation. It has three layers: planning, perception (to find shapes and text), and structural reasoning. These agents work together to transform a simple text prompt into a final, structured graph representation for comparison.

Terminology

Summary

Scientific diagrams convey explicit structural information, yet modern text-to-image models often produce visually plausible but structurally incorrect results. This work introduces SciFlow-Bench, a structure-first benchmark that evaluates diagram generation directly from pixel-level outputs by inverse-parsing generated images back into structured graphs to measure structural recoverability.

SciFlowBench Introduction and Motivation

The paper identifies a visual illusion where image-centric evaluation metrics assign high scores to visually plausible diagrams that contain critical logical errors, such as reversed dependencies or missing components. Existing benchmarks fail to capture these failures because they rely on image-level similarity or intermediate symbolic representations rather than final rendered images. SciFlow-Bench addresses this by adopting a structure-first criterion, evaluating models solely on their ability to preserve the logical relations required for reliable automated or human reconstruction, rather than visual similarity alone.

The SciFlowBench Framework and Evaluation Protocol

SciFlow-Bench is constructed from 500 real-world scientific diagrams, pairing each source framework figure with a canonical ground-truth graph. The evaluation employs a closed-loop, round-trip protocol enabled by a hierarchical multiagent system (HMAS) that coordinates planning, perception, and structural reasoning. This pipeline performs two coupled stages:

  1. Given text descriptions from the source paper, the model generates a diagram image (I).

  2. The generated image is inverse-parsed into predicted graphs to form a unified round-trip protocol for comparison against the canonical ground-truth graph (G∗).

Hierarchical Multi-Agent System (HMAS)

The HMAS serves as the evaluation infrastructure, ensuring consistency across dataset construction and evaluation. It consists of three layers:

  1. Cognitive Planning: The Methodologist extracts method descriptions, and the Visual Translator converts this into a structured visual prompt.

  2. Fine-Grained Perception: Agents like the Environment Curator, Shape Hunter, and Text Spotter operate concurrently to generate perceptual outputs that are fused by the Fusion Arbiter to yield grounded nodes.

  3. Structural Reasoning: The Topology Coder emits a symbolic intermediate representation in Mermaid, which is then parsed by the Graph Architect to yield a structured graph (Gˆ).

Structure-Aware Evaluation Metrics

Performance is evaluated by comparing Gˆ and G∗ in structured graph space using three categories of metrics:

  1. Graph-Level Evaluation: This assesses topology recovery. Node matching relies on semantic similarity of textual descriptions, and edge evaluation focuses on directed dependencies, penalizing missing, reversed, or unsupported connections. The weighting prioritizes relational accuracy; edge-level F1 contributes sixty percent to the graph-level score.

  2. Text-Level Evaluation: This measures consistency between the predicted graph and the structured visual prompt using coverage, faithfulness (penalizing hallucinated elements), and alignment (using GPT-4o).

  3. Image-Level Evaluation: This captures properties not explained by topology alone, including semantic consistency via CLIP similarity, perceptual similarity using LPIPS, and visual flow consistency related to arrow directionality.

Experimental Findings on Model Performance

Experiments reveal a pronounced decoupling between visual fidelity and structural reasoning. Pure diffusion-based generators like SDXL show consistently weak structural recoverability, while autoregressive vision–language models, specifically Gemini 3 Pro Image, exhibit the highest overall scores. The paper demonstrates that structural correctness remains a fundamental challenge, particularly for diagrams with complex topology. Furthermore, ablation studies confirm that components like the Shape Hunter and Text Spotter are essential for achieving balanced performance across metrics. The results suggest that structurally complex diagrams are often accompanied by richer textual descriptions, which top-tier vision–language models can exploit to disambiguate relations more effectively.

Conclusion and Implications

SciFlow-Bench establishes structural recoverability as the central criterion for diagram quality, providing a principled and scalable framework for diagnosing structural failures that are invisible to image-centric metrics. The study concludes that visual plausibility alone does not guarantee structural recoverability, motivating a structure-first evaluation axis necessary for developing multimodal systems that can reason about scientific structure.

--- The gist

SciFlowBench evaluates diagram generation directly from pixel-level outputs by inverse-parsing generated images back into structured graphs to measure structural recoverability, revealing a decoupling between visual fidelity and logical correctness in scientific diagram synthesis.

How it works

SciFlowBench is constructed from 500 real-world scientific diagrams, pairing each source framework figure with a canonical ground-truth graph. The evaluation employs a closed-loop, round-trip protocol enabled by a hierarchical multiagent system (HMAS) that coordinates planning, perception, and structural reasoning. This pipeline performs two coupled stages:

  1. Given text descriptions from the source paper, the model generates a diagram image (I).

Improvements for AI systems

Here are specific improvements for AI systems based on the SciFlow-Bench framework, detailing what those improved systems can achieve:


  1. The fundamental shift from image-centric evaluation (like FID or CLIP Score) to a structure-first, round-trip protocol (inverse parsing into graphs) enables the development of models optimized for logical fidelity rather than just visual realism.

  2. AI systems can be trained with explicit structural supervision derived from canonical ground-truth graphs, ensuring that generated diagrams adhere to predefined scientific workflows and dependencies.

  3. The hierarchical multi-agent system (Cognitive Planning, Perception, Structural Reasoning) provides a modular pipeline for complex diagram generation tasks:

  4. Improved generation systems can utilize the Structured Visual Prompt generated by the Cognitive Planning layer to guide diffusion models more effectively, leading to diagrams with correct component placement and relationship modeling.

  5. Systems can be designed with specialized perception modules (Environment Curator, Shape Hunter, Text Spotter) that focus on recovering functional regions and text accurately, reducing hallucinated components during synthesis.

  6. The Topology Coder agent allows models to explicitly generate a symbolic intermediate representation (like Mermaid code), enabling the system to enforce explicit connectivity decisions and penalize unsupported or reversed dependencies during the generation process.

  7. AI systems can be optimized for robustness against structural complexity, as demonstrated by the performance trends showing that autoregressive vision-language models (like Gemini 3 Pro Image) maintain better structural recoverability on hard, complex topologies compared to vanilla diffusion models (SDXL).

  8. The integration of Text-Level and Image-Level metrics allows for a holistic quality assessment, ensuring the generated diagram is not only structurally sound but also semantically faithful to the source text prompt and visually consistent at the pixel level.

  9. AI systems can be trained using a structure-aware loss function that heavily weights graph-level accuracy (especially edge-level F1 scores) during training, directly addressing the observed decoupling between visual fidelity and structural reasoning.

  10. Systems can incorporate a Minimal Intervention Protocol in their internal verification stages, where components are excluded or added based on explicit textual/visual evidence against the ground truth, leading to higher precision in automated parsing and refinement tools.

This improved AI system can:

  • Generate scientific diagrams that are guaranteed to preserve logical dependencies (e.g., correct arrow directions, missing steps).

  • Produce diagrams that are structurally consistent across varying topological complexities.

  • Be evaluated deterministically using structural metrics rather than subjective visual similarity scores, providing a reliable diagnostic tool for identifying exactly where the model fails in reasoning about scientific structure.

  • Function as a robust evaluation framework capable of distinguishing between visually plausible but structurally flawed outputs and genuinely correct diagrams.

Sources

Related papers