PlotPick: AI-powered batch extraction of numerical data from scientific figures
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "PlotPick: AI-powered batch extraction of numerical data from scientific figures".
Jane: Systematic reviews and meta-analyses frequently require numerical data that authors report only as figures, yet manual digitization is slow and does not scale.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we're looking at this paper today titled "PlotPick: AI-powered batch extraction of numerical data from scientific figures," and the authors are Tommy Carstensen from the Copenhagen Research Centre for Biological and Precision Psychiatry. What’s actually interesting is that they’ve built an open-source tool that uses vision-language models to automatically pull structured tables right out of scientific images, which sounds like it could save so much time.
Jane: It sounds really practical, Tom, especially since manual digitizing data from figures in systematic reviews can be a real bottleneck; the idea is to let an AI handle that tedious part. The paper also tests several different vision-language models against a dedicated chart-to-table model called DePlot to see which ones actually perform better.
Lu: From an AI research standpoint, it’s compelling because it tackles the problem of generalization; dedicated models often struggle with the variety of charts you see in real literature, and this work shows that general-purpose VLMs are surprisingly effective across different chart types.
Meng: It’s good to see a tool that moves beyond just a single use case; if this works well for diverse chart types, it opens up possibilities for automating data extraction across many different fields, which is what we're looking at here at the startup.
Lalam: I think the most significant implication is how this advances our ability to turn visual information into structured knowledge, essentially making complex scientific figures instantly usable in meta-analyses.
Tom: Exactly! It’s not just about extracting numbers; it’s about taking unstructured visual data and turning it into clean, editable tables for researchers who need them fast. So, we're talking about a way to streamline the review process significantly.
Jane: And the paper points out that this tool is open-source, which means anyone can build on this foundation without needing to spend a ton on custom model training or annotation efforts themselves.
Lu: That accessibility is huge; it suggests that powerful extraction capabilities don't have to be locked behind proprietary, highly specialized systems.
Meng: From an engineering viewpoint, the fact that they used PyMuPDF for parsing and then anchored crops based on figure captions shows a solid pipeline approach to handling complex document layouts.
Lalam: And I think the simple prompt they used to ask the VLM for a tab-separated table is pretty impressive; it suggests that general models can handle this kind of structured instruction quite well.
The paper's summary: Tom: Now, let's talk about what PlotPick actually summarizes about the research; they describe it as an open-source tool that takes PDFs, finds figures using text anchors like "Figure one" crops the relevant image and graphics, and then sends that figure to a vision-language model with a very specific instruction to extract the data as a tab-separated table <ref:2605.06021#pg0,an open-source tool that>.
Jane: So, in plain terms, it’s an automated workflow designed for systematic reviews where authors report data only in figures, which usually means you have to manually type everything out. The paper shows this tool batch-extracting numerical data from these scientific figures efficiently.
Lu: What really stands out is the comparison they ran against DePlot and other dedicated models; the results clearly indicate that general-purpose VLMs outperform those specialized tools on both established benchmarks, ChartX and PlotQA, as detailed in page zero of that work.
Meng: The summary points out that these VLMs achieved recall scores between eighty-eight percent and ninety-six percent on ChartX for certain charts, contrasting with DePlot's lower score of seventy-one percent <ref:2605.06021#pg0>. That kind of performance difference is substantial when you’re dealing with large datasets for a review.
Lalam: I find the specific mention that the largest gap in performance is seen on box plots, where DePlot only achieved twenty-four percent RMSF1 while VLMs reached eighty-three percent to ninety-seven percent, really highlights where these general models show their strength <ref:2605.06021#pg0>.
Tom: That contrast is what makes it compelling; it shows that the variety of chart types a general model was trained on gives it an advantage over a model trained only on specific chart distributions. It’s about broad coverage versus narrow specialization.
Jane: The paper emphasizes that this approach requires no model training or manual annotation from the user to get started, which is a major win for researchers who don't have deep machine learning expertise.
Lu: And it addresses the limitation mentioned on page one where dedicated chart-to-table models generalize poorly to diverse charts, suggesting this VLM approach is a necessary step forward for handling real-world scientific documents <ref:2605.06021#pg0>.
Meng: From a practical standpoint, this means researchers could feed in hundreds of figures and get structured data back without needing to hire specialized annotators for every new figure format.
Lalam: It really shows how the capability of these general models, when paired with a smart pipeline like PlotPick, can translate directly into saving significant time and effort in the research lifecycle.
The paper's improvements: Tom: The authors suggest several design choices that make PlotPick effective; they highlight being backend-agnostic, meaning you can swap out different VLMs without needing to rewrite complex prompt engineering, and it supports batch processing for multiple figures in one go.
Jane: And they also talk about the export options available, which include formats like Excel, CSV, LaTeX, JSON, and even R scripts—so the output is directly usable in almost any statistical software or analysis environment.
Lu: The system’s ability to anchor crops based on text blocks matching figure-caption regular expressions is a smart way to ensure that the AI focuses only on the relevant visual data and doesn't get confused by surrounding text or other elements on the page.
Meng: From an engineering angle, the fact that they implemented this batch processing capability means you aren't doing one figure at a time; you can process entire batches efficiently within a single session.
Lalam: The confidence scoring mechanism for each row of extracted data is also a key improvement because it lets downstream users know exactly how reliable each piece of data point actually is, which adds a layer of necessary rigor to the output.
Tom: That confidence score is huge; it gives researchers immediate insight into the uncertainty, allowing them to prioritize their review based on the model's own assessment of its accuracy.
Jane: It’s also interesting how they noted that even with simple prompts, results are still quite close to what you would get from very detailed, manually engineered prompts, which makes it accessible for more people.
Lu: The paper also flags some limitations related to extraction accuracy based on the chart type and model size; specifically, precision is less reliable when the AI has to read axis scales and magnitude with smaller models.
Meng: And they found that a simple two times upscaling of the image preprocessing actually helped improve accuracy by about three percentage points, which is a concrete engineering tweak they found that makes sense <ref:2605.06021#pg1>.
Lalam: So, it’s not just about the extraction itself; it’s about building a robust system around it with good preprocessing and reliable feedback loops.
Conclusion: Tom: To wrap up this discussion on "PlotPick: AI-powered batch extraction of numerical data from scientific figures," we see that general-purpose VLMs demonstrate superior performance compared to dedicated chart-to-table models across the tests, especially when those models are tested on charts they weren't trained on.
Jane: So, the main implication is that for systematic reviews and meta-analyses, this tool offers a scalable way to get structured data from figures without needing massive manual annotation efforts.
Lu: The combination of automatic figure detection and VLM-based extraction makes this a viable method for researchers looking to speed up their data synthesis significantly.
Meng: It’s a practical system that moves the process from being slow and manual to being automated, which is something I’ve seen in many other areas where we can see efficiency gains.
Lalam: I think the confidence scoring and the overall design make this a powerful addition to how we handle visual data in scientific literature today.
Copenhagen Research Centre for Biological and Precision Psychiatry · Mental Health Centre Copenhagen · University Hospital
cs.CV, cs.DL
Submitted: 2026-05-07
Updated: 2026-10-06
Code: https://github.com/tommycarstensen/plotpick
Importance score: 90/100
The gist: Systematic reviews and meta-analyses frequently require numerical data that authors report only as figures, yet manual digitization is slow and does not scale.
Key concepts
- Vision-Language Models (VLMs)
- These are AI models that can understand both images and text. PlotPick uses them to look at a scientific figure, understand what the chart represents, and extract the numbers within it. They are general-purpose tools capable of handling diverse visual data.
- Chart-to-Table Models
- These are specialized AI models trained specifically to convert charts into structured data tables. The paper compares these specialized models against general VLMs and finds that the general VLMs perform better, especially on chart types they weren't specifically trained for.
- Recall (in evaluation)
- This metric measures how many of the actual numbers in the original figure were correctly recovered by the AI. A high recall means the model successfully found a large fraction of all the correct data points, even if some are slightly off.
- RMSF1 (in evaluation)
- This metric is a balanced measure combining precision and recall. It assesses how similar the extracted table is to the true data, considering both accuracy and how well it matches the original structure of the numbers.
Terminology
Summary
Systematic reviews and meta-analyses frequently require numerical data that authors report only as figures, yet manual digitization is slow and does not scale. PlotPick presents an open-source tool that uses vision-language models (VLMs) to batch-extract structured tabular data from scientific figures, demonstrating that general-purpose VLMs outperform dedicated chart-to-table models on diverse chart types.
How it works
PlotPick is a single-file Streamlit application designed for the batch extraction of numerical data from scientific figures. The process begins by uploading PDFs, which are then parsed using PyMuPDF to yield text blocks along with their bounding boxes. Blocks matching a figure-caption regular expression (e.g., “Figure 1”) serve as anchors to crop the region, expanding it to enclose nearby raster images and vector graphics on the same page. Each resulting figure image is subsequently sent to a VLM with a simple extraction prompt: “Extract the data from this chart as a tab-separated table.” The results are then displayed as editable tables, each including per-row confidence scores.
Model Evaluation and Benchmarks
The authors evaluate six VLMs from three different providers—Anthropic (Claude Haiku 4.5, Claude Sonnet 4.6), Google (Gemini 3.1 Flash Lite, Gemini 3 Flash), and OpenAI (GPT-5.4 nano, GPT-5.4 mini)—against a dedicated chart-to-table model named DePlot across two established benchmarks: ChartX and PlotQA. The metrics used for evaluation include Recall, which measures the fraction of ground-truth numeric values recovered with a 5% relative tolerance, and RMSF1 (Relative Mapping Similarity F1), which is the harmonic mean of precision and recall under the same tolerance using permutation-invariant matching.
Performance Against Dedicated Models
The results consistently show that all six VLMs outperform DePlot on both benchmarks. On ChartX (restricted to bar charts, line charts, box plots, and histograms; n=300), VLMs achieve 88–96% recall versus 71% for DePlot. On PlotQA (n=529), the gap is even more pronounced: VLMs achieve 86–99% RMSF1 versus 94% for DePlot. The largest performance gap is observed on chart types absent from the dedicated models’ training data, specifically box plots, where DePlot achieves only 24% RMSF1 while VLMs achieve 83–97%.
Key Design Choices and Findings
Several key design choices underpin the tool's effectiveness. First, PlotPick is Backend-agnostic,
defaulting to Claude Haiku but allowing substitution of any VLM accepting image input without requiring manual prompt engineering, noting that a minimal prompt achieves results within 3 percentage points of a detailed one. Second, the system supports Batch processing,
allowing multiple figures to be processed in a single session with accumulated results. Third, it offers Multiple export formats,
including Excel, CSV, LaTeX, JSON, and R scripts. Furthermore, the discussion indicates that general-purpose VLMs outperform dedicated models due to factors such as Training data scale
and superior Chart-type coverage.
Limitations and Conclusion
The study identifies limitations regarding extraction accuracy based on chart type and model size. Extraction precision is less reliable for charts requiring axis-scale reading, with smaller models showing errors in scale and magnitude. Image preprocessing, specifically a 2× upscaling,
was found to improve accuracy by approximately 3 percentage points. In conclusion, PlotPick demonstrates that general-purpose VLMs can accelerate data extraction for systematic reviews and meta-analyses without the need for model training or manual annotation. The tool is available as open source under the MIT licence.
The gist: General-purpose VLMs outperform dedicated chart-to-table models on established benchmarks across all tested chart types, including those absent from specialized models' training data.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on the PlotPick methodology, and what those improved systems could accomplish:
-
Improve data extraction accuracy for diverse scientific literature figures by leveraging general-purpose Vision-Language Models (VLMs) instead of specialized chart-to-table models.
-
Enable batch processing of multiple complex scientific figures (e.g., bar charts, line charts, and box plots) within a single session without requiring manual annotation or retraining for each chart type encountered in a synthesis project.
-
Develop an AI system capable of converting unstructured scientific PDF documents into structured, editable tabular data (Excel, CSV, LaTeX) suitable for direct input into statistical software or meta-analysis frameworks.
-
Create a robust data extraction pipeline that automatically anchors figure regions based on text captions (e.g.,
Figure 1
) and expands the crop area to include associated vector graphics and raster images, ensuring comprehensive data capture from complex layouts. -
Enhance the reliability of extracted numerical values by incorporating a confidence scoring mechanism for each row of extracted data, allowing downstream researchers to prioritize or flag potentially erroneous extractions based on model uncertainty.
-
Improve the performance of VLM-based extraction systems by utilizing simple, minimal prompts rather than complex, multi-rule prompt engineering, thereby making the system more accessible and less dependent on extensive manual prompt tuning for different chart types.
Sources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models