VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models".
Jane: The paper was written by Qijia He, Xunmei Liu, Hammaad Memon, Ziang Li, Zixian Ma et al. from University of Washington and Allen Institute for Artificial Intelligence and UNC-Chapel Hill.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. We're digging into a new paper today, and it's called "VFig: Vectorizing Complex Figures in SVG with Vision-Language Models." Jane, when you first saw that title, what jumped out at you?
Jane: Tom, the word "complex" is doing so much heavy lifting there. We're not talking about turning a simple smiley face into code. These are scientific diagrams—think of those dense architecture charts you see in AI papers, with boxes, arrows, and text all tangled together.
Tom: Right, and the "SVG" part is the magic. That's the format that lets you scale a diagram without it getting blurry, and it's editable, which is huge. But the problem is, most of the time you only have the flat image, like a PNG.
Jane: Exactly. The original vector file is lost, so you're stuck with a picture you can't tweak. Manually redrawing that? That's a nightmare. This paper is trying to get a computer to do it automatically.
Tom: And they're not just doing it with a simple trick. They built a whole new dataset, a new training method, and a new way to test it. I'm curious about the team behind it, too. It's from the University of Washington and the Allen Institute, with a bunch of collaborators.
Jane: Yeah, and they're clearly thinking about this as a real engineering problem. They're not just showing a cool demo; they're building the infrastructure to make this work reliably. That's what gets me excited.
Tom: For sure. So the big question is, how do you teach a model to look at a messy, complex figure and write clean, structured code that recreates it? That's the puzzle we're going to unpack.
Jane: And I think the answer they came up with is pretty clever. It involves teaching the model to walk before it can run, which we'll get into. But first, let's just appreciate the scale of the problem they're tackling.
Tom: Absolutely. It's one thing to vectorize a logo, but a full research diagram with multiple panels and tiny labels? That's a whole different beast. Let's get into the details.
Summary: Tom: So, Jane, we've set the stage with the title. Now, let's talk about what the paper actually does. The core idea is to train a vision-language model to take a raster image and output SVG code.
Jane: Right, and the first thing they had to do was create the data to train it on. They built something called VFig-Data, which has sixty-six thousand pairs of images and their corresponding SVG code. That's a massive amount of training material.
Tom: And they didn't just scrape random images. They had a whole pipeline. They took real figures from scientific papers and used a powerful model to first describe the figure in detail, then generate the SVG from that description. It's like a two-step process to get higher quality results.
Jane: But they also realized that real-world data is messy. So they generated a separate set of diagrams programmatically, with precise control over shapes, arrows, and colors. That gives them clean, perfect examples to teach the model the basics.
Tom: Then they filtered everything rigorously. They threw out figures that were mostly photos or math equations, because those don't vectorize well. They also filtered out SVG code that was too path-heavy, which would just be a mess of coordinates.
Jane: That filtering is so important. If you train on garbage, you get garbage. They wanted to make sure the model learned to use clean, semantic primitives like rectangles and circles, not just a thousand tiny lines.
Tom: And then they had to actually train the model. They used a two-stage approach. First, they fine-tuned it on simpler diagrams to learn the basics. Then, they fine-tuned it on the complex scientific figures.
Jane: That's the "walk before you run" part. It's like learning to write individual letters before you try to write a full essay. It stabilizes the training and helps the model build a strong foundation.
Tom: But they didn't stop at just supervised learning. They then used reinforcement learning, where the model generates an SVG, renders it, and gets a reward based on how well it matches the original image. That's a really powerful feedback loop.
Jane: And that's where the "vision" part of vision-language models really shines. The model can see its own output and compare it to the target, which is something you can't do with just text.
Tom: So we've got a huge dataset, a clever training curriculum, and a reinforcement learning stage. That's the recipe. But how well does it actually work? That's what we need to look at next.
Jane: Yeah, let's talk about the results and how they measure success. Because "good" is a subjective word when it comes to recreating a diagram.
Improvements: Tom: So, Jane, we've covered the data and the training. Now, the paper introduces a new benchmark called VFig-Bench to see if all that work actually paid off. And it's not just one simple score.
Jane: Right, they knew that a single metric like "does it look similar" isn't enough. A diagram can look similar but have the arrows pointing the wrong way, which would be a disaster. So they created a multi-level evaluation.
Tom: They look at pixel-level similarity, which is the basic stuff. Then they look at component-level scores, like checking if the arrows actually connect the right boxes. And finally, they use a vision-language model as a judge to score the overall structure and details.
Jane: That's clever. It's like grading an essay on grammar, structure, and argument, not just on whether it's spelled correctly. The grammar is the pixels, the structure is the layout, and the argument is whether the connections make sense.
Tom: And the results? Their model, VFig, beats all the other open-source models by a wide margin. It even performs on par with massive proprietary models like GPT-five point two, which is a huge deal for an open-source model.
Jane: It really is. They're showing that with the right data and training strategy, you can close the gap with models that are many times larger. It's not just about scale; it's about being smart about how you use your resources.
Tom: The reinforcement learning stage was key. They found that using a structure-aware reward, which checks for things like connectivity and layout, was much more effective than just using a pixel-level loss.
Jane: That makes sense. Pixel loss might tell you the overall color is right, but it won't tell you that the arrow from box A should point to box B, not box C. The higher-level reward gives the model the right kind of feedback.
Tom: So they're not just making pretty pictures; they're making structurally correct diagrams. That's what makes this so useful for real-world applications.
Jane: And that's the exciting part. This isn't just a research curiosity. This could actually change how people work with scientific figures. Let's think about what that means in practice.
Tom: Absolutely. We've got Lu, Meng, and Lalam here to help us think through the bigger picture. Let's bring them in.
Conclusion: Tom: Okay, let's wrap this up. We've been talking about "VFig: Vectorizing Complex Figures in SVG with Vision-Language Models," and I think we've only scratched the surface of its potential.
Jane: We really have. We've seen how it tackles the hard problem of taking a flat image and turning it into editable, scalable code. And the results are genuinely impressive, matching the big proprietary models.
Tom: For me, the most exciting implication is for accessibility. Think about all the research papers out there with figures that are locked in as images. This could let anyone take that figure, extract the code, and adapt it for their own work.
Jane: That's a great point. It democratizes the ability to edit and reuse scientific visuals. You don't need to be a graphics expert to tweak a diagram for your presentation or your own paper.
Tom: And it's not just for academics. Think about technical documentation, engineering schematics, or even educational materials. Anywhere you have complex diagrams, this could save hours of manual work.
Jane: The key takeaway is that they've shown a clear path forward. By combining a carefully curated dataset, a smart training curriculum, and a structure-aware reward, they've built a model that is both powerful and practical.
Tom: And they're releasing the data and the code, which means other researchers can build on this. That's how the field advances.
Jane: So, as we say goodbye to this paper, we're not just closing a chapter. We're opening the door to a future where the visual language of science is more fluid, more editable, and more accessible to everyone.
Tom: Well said, Jane. That's a perfect note to end on. Thanks to everyone for listening, and we'll see you on the next one.
Qijia He, Xunmei Liu, Hammaad Memon, Ziang Li, Zixian Ma, Jaemin Cho, Zhongzheng Ren, Daniel S Weld, Ranjay Krishna
University of Washington · Allen Institute for Artificial Intelligence · UNC-Chapel Hill
cs.CV, cs.AI
Submitted: 2026-08-16
Updated: 2026-08-18
Code: https://github.com/RAIVNLab/VFig
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 69/100
Key concepts
- SVG
- Scalable Vector Graphics is a format that allows diagrams to be scaled without becoming blurry and remains editable. The paper focuses on converting complex figures from raster images (like PNGs) into this vector format so users can tweak them.
- VFig-Data
- This is the massive training dataset created for the model, containing sixty-six thousand pairs of images and their corresponding SVG code. The data was created using a two-step pipeline involving detailed descriptions and programmatic generation to ensure high quality.
- Reinforcement Learning
- This training method involves a feedback loop where the model generates an SVG, renders it, and receives a reward based on how well it matches the original image. A structure-aware reward is used to guide the model toward correct layout and connectivity.
- VFig-Bench
- This is a new benchmark created to test the model's performance. It uses a multi-level evaluation, checking pixel similarity, component connections, and overall structure using another vision-language model as a judge.
Terminology
Summary
Summary
This paper introduces VFig, a family of Vision-Language Models (VLMs) trained for the task of converting complex rasterized figures (e.g., PNG or JPEG) into high-fidelity, editable Scalable Vector Graphics (SVG) code. The authors motivate the work by noting that original vector source files for technical illustrations and scientific diagrams are frequently lost, leaving only flat rasterized versions that are difficult to modify or scale, and that manual reconstruction is prohibitively labor-intensive.
The paper makes four main contributions: a large-scale dataset, a two-stage training strategy, a comprehensive evaluation suite, and a systematic empirical study.
1. Data Contribution (VFig-Data): The authors construct VFig-Data, a large-scale dataset of 66K high-quality figure–SVG pairs. It contains two complementary subsets:
-
VFig-Data-Complex-Diagrams: Real-world scientific paper figures collected as raster images and converted into structured SVG using a two-stage pipeline. First, a VLM (Gemini-3-Pro) produces a structured description of the figure, capturing geometric elements, text, spatial layout, and relationships. Second, the VLM generates SVG code conditioned on both the original image and the description. This pipeline was preferred in 88.7% of pairwise comparisons in a human study.
-
VFig-Data-Shapes-and-Arrows: Programmatically generated diagrams with precise control over visual attributes (shapes, arrows, fonts, styles), synthesized directly in SVG using 19 layout templates with randomized parameters.
The data pipeline includes rigorous filtering: Image Filtering removes figures dominated by natural images, screenshots, math equations, plots, and tables, retaining only diagram-centric figures. Code Filtering removes SVG outputs dominated by free-form elements, retaining figures composed of semantically meaningful primitives (e.g.,, , ). The authors also incorporate 78K data points from academic datasets (SVG-Diagrams and Molmo2-Diagram) after similar filtering. The training mixture statistics are summarized in Table 1, with VFig-Data-Complex-Diagrams showing the highest structural complexity (55.3) and element complexity (4.0).
2. Training Contribution (Two-Stage Strategy): The authors propose a coarse-to-fine training curriculum:
-
Supervised Fine-Tuning (SFT): The model is first trained on structurally simpler diagrams (SVG-Diagrams, Molmo2-Diagram, VFig-Data-Shapes-and-Arrows) to establish robust primitive-level generation and basic layout understanding. It is then fine-tuned on complex scientific figures (VFig-Data-Complex-Diagrams) to develop compositional reasoning and structural fidelity. The SFT objective maximizes the likelihood of training data: L SFT = -E[log p θ(yx)].
-
Reinforcement Learning (RL) with Visual Feedback: To close the gap between token-level likelihood and visual quality, the authors apply Group Relative Policy Optimization (GRPO). For each input figure, the model samples multiple SVG programs, which are rendered and scored by a reward function. The reward is an unweighted average of four rubric scores from a VLM judge (Gemini-3-Flash): Presence (all required visual elements present), Layout (spatial arrangement and alignment), Connectivity (arrows and lines connect correct endpoints), and Details (text accuracy and fine styling). The GRPO objective includes KL regularization against the SFT checkpoint. The authors note that the VLM judge scores exhibit strong Pearson correlation with human judgments (overall r = 0.89).
3. Evaluation Contribution (VFig-Bench): The authors introduce VFig-Bench, a benchmark with 392 realistic scientific figures held out from VFig-Data-Complex-Diagrams. It features a coarse-to-fine evaluation protocol with three granularities: pixel-level metrics (SSIM, LPIPS, VisualSim), component-level scores (rule-based arrow and shape matching), and image-level judgments (VLM judges from Gemini and GPT). They also report SVG cleanliness and render rate. The rule-based evaluation is detailed in the appendix, with shape attributes (label, type, fill color/style, stroke color/style, position, font, aspect ratio) and arrow attributes (source/destination, head, head size, curve, color, overlap).
4. Experimental Findings: The paper answers four research questions:
-
RQ1 (Current VLM capability): Classical raster-to-vector methods (VTracer) achieve high pixel similarity but fail to generate clean primitives. Open-source VLM baselines perform worse in both visual fidelity and structural correctness. VFig achieves state-of-the-art performance among open-source models and performs on par with GPT-5.2, achieving a VLM-Judge score of 0.829 on VFig-Bench.
-
RQ2 (Curriculum SFT): The two-stage curriculum improves render success rate significantly (e.g., from 0.749 to 0.933 for Qwen3-VL-4B) and slightly increases VLM-judge scores.
-
RQ3 (RL with visual feedback): RL consistently improves generation quality over SFT alone across all metrics and datasets (e.g., VLM-Judge improves from 0.712 to 0.804 on average).
-
RQ4 (Reward granularity): Structure-aware VLM-based rewards are more effective than pixel-level objectives. Removing any of the four rubric components degrades judge-based metrics, with the largest drops from removing layout or details. Adding pixel-level objectives (Gemini + Pixel) slightly improves SSIM but reduces judge-based scores.
Ablations: The paper reports ablations on backbone choice (Qwen3-VL outperforms InternVL3.5 and Qwen2.5-VL), LoRA rank (rank 64 is best), SFT target modules (LM-only is best), RL initialization (two-stage SFT is better), and model size (8B is slightly better than 4B but with a trade-off in code cleanliness).
Human Evaluation: A human evaluation with pairwise comparisons confirms the benchmark results. Gemini 3 Pro achieves the highest Elo rating (1852.5), followed by GPT 5.2 (1617.3), VFig (1473.8), and Qwen3-VL-4B (1056.4). VFig wins 81.6% of decisive comparisons against Qwen3-VL-4B.
Limitations: The paper acknowledges failure cases, primarily in fine-grained geometric details (thin lines, arrows, small text), exact colors, and subtle stylistic details. Diagrams with 3D shapes or perspective-like objects are especially challenging. The authors suggest that further improvements may require richer reward design and broader training data coverage beyond structured scientific diagrams.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement and the resulting capabilities of the improved AI system:
-
Implementation: Replace single-stage training with a progressive curriculum. Stage 1 trains on simple diagrams (icons, basic shapes, arrows) to establish primitive-level generation. Stage 2 fine-tunes on complex scientific figures (multi-panel, hierarchical compositions).
-
Result: The model learns atomic SVG primitives (rect, circle, text) before tackling compositional reasoning, improving render success rate from 74.9% to 93.3% and VLM-Judge score from 0.712 to 0.737.
-
Implementation: After SFT, apply GRPO with a VLM judge (Gemini-3-Flash) that scores outputs across four dimensions: presence, layout, connectivity, and details. The reward is the unweighted average of these four scores, with invalid renders receiving zero reward.
-
Result: Improves LPIPS from 0.264 to 0.212, VisualSim from 0.951 to 0.957, and VLM-Judge from 0.788 to 0.829 compared to SFT alone.
-
Implementation: Filter training data to retain SVGs where basic shapes + connectors comprise ≥40% of geometric elements and complex shapes (path, polygon) ≤50. This removes path-heavy tracing outputs that cause token explosion.
-
Result: Produces cleaner, more editable SVG code with higher semantic cleanliness scores (0.842 vs 0.761 for baseline), improving downstream editability.
-
Implementation: First prompt a VLM to generate a structured textual description of the figure (capturing geometry, text, layout, relationships), then prompt it again to generate SVG code conditioned on both the image and description.
-
Result: Improves layout accuracy, text rendering, and shape selection by 88.7% in human preference compared to single-pass generation.
-
Implementation: Evaluate at three levels: pixel-level (SSIM, LPIPS), component-level (rule-based shape/arrow matching), and image-level (VLM judge scores). This provides a comprehensive view beyond single-metric evaluation.
-
Result: Enables detection of structural errors (e.g., wrong arrow connections) that pixel metrics miss, aligning better with human perception (Pearson r=0.89).
-
Convert complex scientific figures to editable SVG code with 93.3% render success rate, preserving hierarchical structure, precise alignments, and connectivity.
-
Generate SVG code that is 96.4% visually similar to the input raster image (VisualSim score), with structural correctness (presence, layout, connectivity, details) scoring 0.829/1.0.
-
Handle multi-panel diagrams, flowcharts, and architecture diagrams with nested layouts, heterogeneous primitives, and intricate connectivity that previous models failed to reconstruct.
-
Produce semantically clean SVG code (0.842 cleanliness score) using primitives like rect, circle, and line rather than path-heavy tracing, making outputs editable and reusable.
-
Achieve performance on par with GPT-5.2 (VLM-Judge 0.829 vs 0.828) while being open-source and 4B parameters, demonstrating that targeted data curation and structured training can close the gap with much larger proprietary models.
-
Maintain structural integrity in edge cases: correct arrow endpoints, proper shape grouping, and accurate text placement, even in figures with 50+ elements and complex interconnections.
Abstract
Scalable Vector Graphics (SVG) are an essential format for technical illustration and digital design, offering precise resolution independence and flexible semantic editability. In practice, however, original vector source files are frequently lost or inaccessible, leaving only "flat" rasterized versions (e.g., PNG or JPEG) that are difficult to modify or scale. Manually reconstructing these figures is a prohibitively labor-intensive process, requiring specialized expertise to recover the original geometric intent. To bridge this gap, we propose VFIG, a family of Vision-Language Models trained for complex and high-fidelity figure-to-SVG conversion. While this task is inherently data-driven, existing datasets are typically small-scale and lack the complexity of professional diagrams. We address this by introducing VFIG-DATA, a large-scale dataset of 66K high-quality figure-SVG pairs, curated from a diverse mix of real-world paper figures and procedurally generated diagrams. Recognizing that SVGs are composed of recurring primitives and hierarchical local structures, we introduce a coarse-to-fine training curriculum that begins with supervised fine-tuning (SFT) to learn atomic primitives and transitions to reinforcement learning (RL) refinement to optimize global diagram fidelity, layout consistency, and topological edge cases. Finally, we introduce VFIG-BENCH, a comprehensive evaluation suite with novel metrics designed to measure the structural integrity of complex figures. VFIG achieves state-of-the-art performance among open-source models and performs on par with GPT-5.2, achieving a VLM-Judge score of 0.829 on VFIG-BENCH.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- DeepSVG: A Hierarchical Generative Network for Vector Graphics Animation
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- FIGR: Few-shot Image Generation with Reptile
- A Neural Representation of Sketch Drawings
- LoRA: Low-Rank Adaptation of Large Language Models
- Towards Layer-wise Image Vectorization
- DINOv2: Learning Robust Visual Features without Supervision
- Learning Transferable Visual Models From Natural Language Supervision
- Im2Vec: Synthesizing Vector Graphics without Vector Supervision
- StarVector: Generating Scalable Vector Graphics Code from Images and Text
- Rendering-Aware Reinforcement Learning for Vector Graphics Generation
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- DeepVecFont: Synthesizing High-quality Vector Fonts via Dual-modality Learning
- Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning
- Empowering LLMs to Understand and Generate Complex Vector Graphics
- OmniSVG: A Unified Scalable Vector Graphics Generation Model
- Sigmoid Loss for Language Image Pre-Training
- Beyond Pixels: Exploring Human-Readable SVG Generation for Simple Images with Vision Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models