VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models

summary

Video file (mp4)

In short

The episode discusses the paper "VFig: Vectorizing Complex Figures in SVG with Vision-Language Models," which trains a vision-language model to convert raster images of complex scientific figures into editable SVG code. The hosts detail the massive dataset creation, two-stage training process using reinforcement learning, and a multi-level evaluation benchmark called VFig-Bench. They conclude that this technology could democratize editing scientific visuals.

Key concepts

SVG
Scalable Vector Graphics is a format that allows diagrams to be scaled without becoming blurry and remains editable. The paper focuses on converting complex figures from raster images (like PNGs) into this vector format so users can tweak them.
VFig-Data
This is the massive training dataset created for the model, containing sixty-six thousand pairs of images and their corresponding SVG code. The data was created using a two-step pipeline involving detailed descriptions and programmatic generation to ensure high quality.
Reinforcement Learning
This training method involves a feedback loop where the model generates an SVG, renders it, and receives a reward based on how well it matches the original image. A structure-aware reward is used to guide the model toward correct layout and connectivity.
VFig-Bench
This is a new benchmark created to test the model's performance. It uses a multi-level evaluation, checking pixel similarity, component connections, and overall structure using another vision-language model as a judge.

Terminology used across episodes

This episode discusses

The paper

VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models · Read on arXiv

Qijia He, Xunmei Liu, Hammaad Memon, Ziang Li, Zixian Ma, Jaemin Cho, Zhongzheng Ren, Daniel S Weld, Ranjay Krishna

University of Washington · Allen Institute for Artificial Intelligence · UNC-Chapel Hill

Scalable Vector Graphics (SVG) are an essential format for technical illustration and digital design, offering precise resolution independence and flexible semantic editability. In practice, however, original vector source files are frequently lost or inaccessible, leaving only "flat" rasterized versions (e.g., PNG or JPEG) that are difficult to modify or scale. Manually reconstructing these figures is a prohibitively labor-intensive process, requiring specialized expertise to recover the original geometric intent. To bridge this gap, we propose VFIG, a family of Vision-Language Models trained for complex and high-fidelity figure-to-SVG conversion. While this task is inherently data-driven, existing datasets are typically small-scale and lack the complexity of professional diagrams. We address this by introducing VFIG-DATA, a large-scale dataset of 66K high-quality figure-SVG pairs, curated from a diverse mix of real-world paper figures and procedurally generated diagrams. Recognizing that SVGs are composed of recurring primitives and hierarchical local structures, we introduce a coarse-to-fine training curriculum that begins with supervised fine-tuning (SFT) to learn atomic primitives and transitions to reinforcement learning (RL) refinement to optimize global diagram fidelity, layout consistency, and topological edge cases. Finally, we introduce VFIG-BENCH, a comprehensive evaluation suite with novel metrics designed to measure the structural integrity of complex figures. VFIG achieves state-of-the-art performance among open-source models and performs on par with GPT-5.2, achieving a VLM-Judge score of 0.829 on VFIG-BENCH.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models".

Jane: The paper was written by Qijia He, Xunmei Liu, Hammaad Memon, Ziang Li, Zixian Ma et al. from University of Washington and Allen Institute for Artificial Intelligence and UNC-Chapel Hill.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. We're digging into a new paper today, and it's called "VFig: Vectorizing Complex Figures in SVG with Vision-Language Models." Jane, when you first saw that title, what jumped out at you?

Jane: Tom, the word "complex" is doing so much heavy lifting there. We're not talking about turning a simple smiley face into code. These are scientific diagrams—think of those dense architecture charts you see in AI papers, with boxes, arrows, and text all tangled together.

Tom: Right, and the "SVG" part is the magic. That's the format that lets you scale a diagram without it getting blurry, and it's editable, which is huge. But the problem is, most of the time you only have the flat image, like a PNG.

Jane: Exactly. The original vector file is lost, so you're stuck with a picture you can't tweak. Manually redrawing that? That's a nightmare. This paper is trying to get a computer to do it automatically.

Tom: And they're not just doing it with a simple trick. They built a whole new dataset, a new training method, and a new way to test it. I'm curious about the team behind it, too. It's from the University of Washington and the Allen Institute, with a bunch of collaborators.

Jane: Yeah, and they're clearly thinking about this as a real engineering problem. They're not just showing a cool demo; they're building the infrastructure to make this work reliably. That's what gets me excited.

Tom: For sure. So the big question is, how do you teach a model to look at a messy, complex figure and write clean, structured code that recreates it? That's the puzzle we're going to unpack.

Jane: And I think the answer they came up with is pretty clever. It involves teaching the model to walk before it can run, which we'll get into. But first, let's just appreciate the scale of the problem they're tackling.

Tom: Absolutely. It's one thing to vectorize a logo, but a full research diagram with multiple panels and tiny labels? That's a whole different beast. Let's get into the details.

Summary: Tom: So, Jane, we've set the stage with the title. Now, let's talk about what the paper actually does. The core idea is to train a vision-language model to take a raster image and output SVG code.

Jane: Right, and the first thing they had to do was create the data to train it on. They built something called VFig-Data, which has sixty-six thousand pairs of images and their corresponding SVG code. That's a massive amount of training material.

Tom: And they didn't just scrape random images. They had a whole pipeline. They took real figures from scientific papers and used a powerful model to first describe the figure in detail, then generate the SVG from that description. It's like a two-step process to get higher quality results.

Jane: But they also realized that real-world data is messy. So they generated a separate set of diagrams programmatically, with precise control over shapes, arrows, and colors. That gives them clean, perfect examples to teach the model the basics.

Tom: Then they filtered everything rigorously. They threw out figures that were mostly photos or math equations, because those don't vectorize well. They also filtered out SVG code that was too path-heavy, which would just be a mess of coordinates.

Jane: That filtering is so important. If you train on garbage, you get garbage. They wanted to make sure the model learned to use clean, semantic primitives like rectangles and circles, not just a thousand tiny lines.

Tom: And then they had to actually train the model. They used a two-stage approach. First, they fine-tuned it on simpler diagrams to learn the basics. Then, they fine-tuned it on the complex scientific figures.

Jane: That's the "walk before you run" part. It's like learning to write individual letters before you try to write a full essay. It stabilizes the training and helps the model build a strong foundation.

Tom: But they didn't stop at just supervised learning. They then used reinforcement learning, where the model generates an SVG, renders it, and gets a reward based on how well it matches the original image. That's a really powerful feedback loop.

Jane: And that's where the "vision" part of vision-language models really shines. The model can see its own output and compare it to the target, which is something you can't do with just text.

Tom: So we've got a huge dataset, a clever training curriculum, and a reinforcement learning stage. That's the recipe. But how well does it actually work? That's what we need to look at next.

Jane: Yeah, let's talk about the results and how they measure success. Because "good" is a subjective word when it comes to recreating a diagram.

Improvements: Tom: So, Jane, we've covered the data and the training. Now, the paper introduces a new benchmark called VFig-Bench to see if all that work actually paid off. And it's not just one simple score.

Jane: Right, they knew that a single metric like "does it look similar" isn't enough. A diagram can look similar but have the arrows pointing the wrong way, which would be a disaster. So they created a multi-level evaluation.

Tom: They look at pixel-level similarity, which is the basic stuff. Then they look at component-level scores, like checking if the arrows actually connect the right boxes. And finally, they use a vision-language model as a judge to score the overall structure and details.

Jane: That's clever. It's like grading an essay on grammar, structure, and argument, not just on whether it's spelled correctly. The grammar is the pixels, the structure is the layout, and the argument is whether the connections make sense.

Tom: And the results? Their model, VFig, beats all the other open-source models by a wide margin. It even performs on par with massive proprietary models like GPT-five point two, which is a huge deal for an open-source model.

Jane: It really is. They're showing that with the right data and training strategy, you can close the gap with models that are many times larger. It's not just about scale; it's about being smart about how you use your resources.

Tom: The reinforcement learning stage was key. They found that using a structure-aware reward, which checks for things like connectivity and layout, was much more effective than just using a pixel-level loss.

Jane: That makes sense. Pixel loss might tell you the overall color is right, but it won't tell you that the arrow from box A should point to box B, not box C. The higher-level reward gives the model the right kind of feedback.

Tom: So they're not just making pretty pictures; they're making structurally correct diagrams. That's what makes this so useful for real-world applications.

Jane: And that's the exciting part. This isn't just a research curiosity. This could actually change how people work with scientific figures. Let's think about what that means in practice.

Tom: Absolutely. We've got Lu, Meng, and Lalam here to help us think through the bigger picture. Let's bring them in.

Conclusion: Tom: Okay, let's wrap this up. We've been talking about "VFig: Vectorizing Complex Figures in SVG with Vision-Language Models," and I think we've only scratched the surface of its potential.

Jane: We really have. We've seen how it tackles the hard problem of taking a flat image and turning it into editable, scalable code. And the results are genuinely impressive, matching the big proprietary models.

Tom: For me, the most exciting implication is for accessibility. Think about all the research papers out there with figures that are locked in as images. This could let anyone take that figure, extract the code, and adapt it for their own work.

Jane: That's a great point. It democratizes the ability to edit and reuse scientific visuals. You don't need to be a graphics expert to tweak a diagram for your presentation or your own paper.

Tom: And it's not just for academics. Think about technical documentation, engineering schematics, or even educational materials. Anywhere you have complex diagrams, this could save hours of manual work.

Jane: The key takeaway is that they've shown a clear path forward. By combining a carefully curated dataset, a smart training curriculum, and a structure-aware reward, they've built a model that is both powerful and practical.

Tom: And they're releasing the data and the code, which means other researchers can build on this. That's how the field advances.

Jane: So, as we say goodbye to this paper, we're not just closing a chapter. We're opening the door to a future where the visual language of science is more fluid, more editable, and more accessible to everyone.

Tom: Well said, Jane. That's a perfect note to end on. Thanks to everyone for listening, and we'll see you on the next one.

More episodes

← Home