AnchorFlow: Learning Anchor Placement for Faithful and Editable SVG Reconstruction

arXiv:2605.19551 · cs.GR, cs.CV · Submitted 2026-05-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "AnchorFlow: Learning Anchor Placement for Faithful and Editable SVG Reconstruction".

Tom: The gist AnchorFlow proposes an editable SVG reconstruction framework that models path-level anchor placement with sparse anchor point fields to achieve a favorable fidelity-editability trade-off Anchor Placement as Structural Scaffold…

Jane: First, who's behind it and why it matters.

Paper summary: Tom: Welcome back folks! We're diving into a really interesting paper today called "AnchorFlow: Learning Anchor Placement for Faithful and Editable SVG Reconstruction."

Jane: It tackles this big problem in image-to-SVG reconstruction, which is making vector graphics that look exactly like the original picture but are also easy to actually change later.

Tom: So, the main idea here is addressing that trade-off between getting high fidelity and keeping the vector structure sparse and editable.

Jane: They argue that this structural choice should be made at the level of anchor placement because anchors on Bézier curves define how local path structure works and really affect both accuracy and editability.

Tom: That sounds deep, Jane. So, what exactly is AnchorFlow proposing as their solution?

Lu: It introduces something they call a sparse anchor field. Think of it like this: instead of just drawing the final lines immediately, they learn an intermediate representation that shows where the anchors are *likely* to be, using peaks for likely locations and weaker support for connectivity.

Meng: So it's not the final vector output yet, but a learned structural step that helps guide stable anchor placement before it turns into a proper Bézier path.

Jane: Right, so they build this field first and then use it to create the actual SVG path structure.

Tom: And they lay out this pipeline with three main stages: predicting the anchor field, then using that field to do a hard resolution of the path, and finally adding a rendering-guided refinement stage.

Meng: That refinement step is where things get cool; they feed back errors from how the image actually renders into the system to update that anchor field and fix structural mistakes when necessary.

Jane: It sounds like they're trying to use the visual feedback loop to keep the structure clean during reconstruction.

Lu: They describe this sparse anchor field mathematically using a formula that combines evidence for sharp anchor peaks with support for path ordering and connectivity, which is what keeps things connected even when evidence is weak.

Tom: So it’s a learned way to balance finding the main structural points versus just following every little pixel edge.

Jane: And they show how this approach handles noisy input, specifically when the local path evidence is messy or the boundaries are slightly perturbed.

Meng: That’s important because traditional methods often get stuck with fragmented paths or redundant anchors when the input isn't perfectly clean.

Tom: So, what does this mean for someone who just watches videos or looks at images?

Jane: It means they can reconstruct compact and editable SVGs even from isolated paths or full multi-part images without getting hung up on tiny boundary artifacts.

Tom: They found that the learned anchor field stays stable even when the local path evidence is noisy, which leads to more stable placement of the anchors.

Paper summary: Meng: From an engineering standpoint, I’m interested in how they manage that iterative correction loop; it sounds like a bit slower than just tracing a path in one go because you're constantly refining based on rendering feedback.

Tom: That iterative nature is what lets them correct those local structural errors, but it makes the process take longer than a single pass method.

Lu: They show that this method produces compact editable SVGs with competitive raster fidelity across different benchmarks compared to baselines like AdaVec.

Jane: So, they managed to hit that sweet spot where you get a neat vector file that looks really good and is easy for a designer to work with.

Tom: The authors of "AnchorFlow: Learning Anchor Placement for Faithful and Editable SVG Reconstruction" are Mengnan Jiang, Christian Franke, Michele Franco Adesso, Antonio Haas, and Grace Li Zhang from the Technical University of Darmstadt.

Jane: They’re really pushing the idea that focusing on anchor placement is a better way to solve the fidelity versus editability problem than focusing only on tracing or optimization methods.

Meng: If you look at what they did for full-image reconstruction, they still have to contend with errors inherited from the component extraction frontend, and for really ambiguous or detailed parts, the local paths can still end up imperfect.

Tom: That’s a fair caveat; it isn't perfect across every single complex image.

Jane: Exactly. So while AnchorFlow is very robust against noisy local evidence, it can still inherit issues from earlier steps in the pipeline if the initial component extraction wasn't perfect for that specific image.

Tom: It really shows how much of vectorization depends on where you start the process and what kind of evidence you feed into it. So, this paper is a big step toward making SVG reconstruction practical for real-world messy data.

Jane: And it gives us a framework where the AI learns to build the structure first based on anchors, rather than just guessing every single curve parameter immediately.

Tom: It’s about learning the skeleton before you worry about drawing all the skin on top of it.

Lu: The creative potential here is huge because this sparse field could become a really powerful way for AI to understand and manipulate 2D graphics in a way that respects both visual accuracy and human-friendly editability <ref:2605.19551#pg1>.

Meng: I wonder how robust this approach would be if we tried to apply it to, say, highly stylized or abstract art where the boundaries are inherently less clear than in a simple product icon.

Tom: That’s a good thought for the future work, Lu. It opens up possibilities for more nuanced vector generation beyond clean lines.

Jane: Overall, "AnchorFlow: Learning Anchor Placement for Faithful and Editable SVG Reconstruction" gives us a concrete way to handle that structural trade-off by learning where the anchors should go first.

Tom: And that’s all we have time for today on this paper!

Conclusion: Tom: So we've been looking at AnchorFlow, which is this new way to make SVG files from images that keeps things editable.

Jane: It’s all about learning where the anchor points should be placed before you even draw the final lines.

Lu: It treats anchor placement as a structural scaffold, so it models the path by figuring out those sparse anchor locations first.

Meng: So they’re using this learned field to guide stable anchor placement, which is a clever way to handle the fidelity versus editability trade-off.

Lalam: The core idea is that these anchors define the local path structure and strongly influence both accuracy and how easy it is to change later.

Tom: But what does this actually mean for us when we look at the conclusion of this AnchorFlow paper?

Jane: Well, they show that even with a lot of messy input, AnchorFlow produces compact and editable SVGs that still look pretty good.

Lu: They found that the learned anchor field stays stable even when local path evidence is noisy or the edges are slightly disturbed.

Meng: That stability is key for practical application because it means the reconstruction doesn't fall apart just because one small part of an image was a little unclear.

Tom: The authors are essentially saying they’ve found a method that gets you a good balance between looking real and being able to change the file easily.

Jane: They also note that this approach is preferred over some other methods when it comes to how much you can actually edit the resulting vector file.

Lu: It points toward an architecture where the AI focuses on structural anchors rather than just trying to trace every single pixel immediately.

Meng: That’s interesting because it suggests a more structured way for AI to approach image reconstruction tasks that involve geometry.

Tom: This moves the focus from just drawing lines to learning the underlying structure first, which is a pretty fundamental shift in how we think about this kind of AI work.

Jane: It shows that by modeling anchor placement sparsely, you get a more reliable starting point for generating vector graphics from visual input.

Mercedes-Benz AG · Technical University of Darmstadt

cs.GR, cs.CV

Submitted: 2026-05-19

Updated: 2026-10-08

Comments: 22 pages, including supplementary material. Revised version of the same work; title, method description, and experimental evaluation updated

Code: https://github.com/googlefonts/noto-emoji

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 85/100

The gist: The gist AnchorFlow proposes an editable SVG reconstruction framework that models path-level anchor placement with sparse anchor point fields to achieve a favorable fidelity-editability trade-off

Key concepts

Sparse Anchor Field
This is an intermediate representation that models where anchors (key points on a path) are likely to be placed. It uses sharp peaks for anchor locations and weaker contour support to indicate path connectivity, acting as a learned structural scaffold rather than the final SVG output.
Anchor Flow Pipeline
This is the three-step process: first, predicting the sparse anchor field using AFNet; second, converting that field into an explicit SVG path by selecting and ordering anchors; and third, refining the field using rendering errors to correct structural mistakes.
Rendering-Guided Field Refinement
This stage uses actual rendering feedback to iteratively improve the anchor field. It minimizes a loss function that balances increasing responses in missing areas with suppressing unsupported activations, ensuring the final path structure aligns well with how it looks when drawn.
Fidelity-Editability Trade-off
The core goal is to find a balance between how closely the reconstructed vector matches the original image (fidelity) and how easily a human can modify that vector (editability). AnchorFlow aims to achieve this by focusing on stable structural anchors rather than noisy local details.

Terminology

Summary

The gist AnchorFlow proposes an editable SVG reconstruction framework that models path-level anchor placement with sparse anchor point fields to achieve a favorable fidelity-editability trade-off

Anchor Placement as Structural Scaffold

The central argument is that the structural trade-off in vectorization should be addressed at the level of anchor placement, since anchors on Bézier curves define local path structure and strongly affect both accuracy and editability AnchorFlow models this by introducing a sparse anchor field, an image-conditioned representation whose peaks indicate likely anchor locations and whose local support provides evidence for path connectivity This field is not itself the final vector output but rather a learned structural intermediate that can guide stable anchor placement before being parsed into an ordered Bézier path

AnchorFlow Pipeline

AnchorFlow builds on this representation with three core stages: anchor field prediction, field-conditioned hard resolution, and rendering-guided field refinement

  1. Anchor Field Prediction: A lightweight predictor, Anchor Field Net (AFNet), predicts a sparse anchor field for each local path instance

  2. Field-Conditioned Hard Resolution: This stage converts the predicted field into an explicit SVG path by selecting anchors, ordering them, establishing connectivity, and initializing Bézier segments

  3. Rendering-Guided Field Refinement: This introduces a mechanism where rendering errors are fed back to update the anchor field and re-resolve the path when needed

Sparse Anchor Field Representation

The key design of AnchorFlow is to model anchor placement before constructing the SVG path, introducing a sparse anchor field as an intermediate representation between raster evidence and editable vector structure For a local path instance Xm, the target field is conceptually built from sharp anchor peaks and weaker contour support using the formula F∗m(p) = clip max a∗i ∈A∗m exp −∥p − a∗i∥2 / 2σa2 + λΓ exp − d(p, Γm)2 / 2σΓ2, where the first term provides sparse evidence for anchor locations, while the second term supports path ordering and connectivity

Rendering-Guided Field Refinement

The refinement stage uses true rendering feedback to correct local structural errors by optimizing a bounded perturbation in the latent space The objective function Lrefine minimizes Lrefine = λ+ W+, ψτ (Fm) + λ− W−, Fm + λf∥Fm − Fm,0∥1 + λz∥zm − zm,0∥2 / 2, where the positive guidance term increases field responses in missing regions and the negative guidance term suppresses unsupported activation The candidate is kept only if it improves the tolerance-based stroke score Fδ or satisfies an acceptance threshold τF; otherwise, AnchorFlow keeps the previous path

Robustness and Evaluation

AnchorFlow demonstrates robustness under boundary perturbations, showing that the learned anchor field supports stable anchor placement even when local path evidence is noisy or boundary perturbed In single-path reconstruction, AnchorFlow preserves a compact path structure because the learned anchor field remains stable and focuses on structural anchor locations rather than local boundary fluctuations Furthermore, human evaluation shows that AnchorFlow is preferred over AutoTrace [1], LIVE [17], and AdaVec [31] for editability The full-image reconstruction experiments show that AnchorFlow achieves a favorable fidelity-editability trade-off across benchmarks, consistently using fewer editable parameters than the AdaVec [31] baseline Overall, AnchorFlow produces compact editable SVGs with competitive raster fidelity

Limitations

AnchorFlow has several limitations, including a modest train-inference gap because the predictor is trained with fixed anchor targets but refined with rendering feedback at inference time In full-image reconstruction, it can also inherit errors from the component extraction frontend, and highly ambiguous or detailed components may still produce imperfect local paths The iterative correction loop is slower than single-pass tracing methods, motivating more efficient structure-aware refinement

Conclusion

AnchorFlow presents an editable SVG reconstruction framework based on sparse anchor point fields that recovers compact and input-faithful SVGs from both isolated paths and full images AnchorFlow has several limitations, including a modest train-inference gap because the predictor is trained with fixed anchor targets but refined with rendering feedback at inference time In full-image reconstruction, it can also inherit errors from the component extraction frontend, and highly ambiguous or detailed components may still produce imperfect local paths The iterative correction loop is slower than single-pass tracing methods, motivating more efficient structure-aware refinement

References

[1] AutoTrace Project. AutoTrace: Bitmap to vector graphics converter. https://autotrace.sourceforge.net/, 2024 >

[2] Alexandre Carlier, Martin Danelljan, Alexandre Alahi, and Radu Timofte. DeepSVG: A hierarchical generative network for vector graphics animation. In Advances in Neural Information Processing Systems, volume 33, pages 16351–16361, 2020 >

[3] Souymodip Chakraborty, Vineet Batra, Ankit Phogat, Vishwas Jain, Jaswant Singh Ranawat, Sumit Dhingra, Kevin Wampler, and Michal Lukác. Image vectorization via gradient reconstruction. ˇ Computer Graphics Forum, 44(2):e70055, 2025 >

[4] Zehao Chen and Rong Pan. Svgbuilder: Component-based colored svg generation with text-guided autoregressive transformers. In Proceedings of the AAAI Conference on Artificial Intelligence, 2025 >

[5] Ayan Das, Yongxin Yang, Timothy Hospedales, Tao Xiang, and Yi-Zhe Song. Cloud2curve: Generation and vectorization of parametric sketches. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7088–7097, 2021 >

[6] Maria Dziuba, Ivan Jarsky, Valeria Efimova, and Andrey Filchenkov. Image vectorization: A review. arXiv preprint arXiv:2306.06441, 2023 >

[7] Vage Egiazarian, Oleg Voynov, Alexey Artemov, Denis Volkhonskiy, Aleksandr Safin, Maria Taktasheva, Denis Zorin, and Evgeny Burnaev. Deep vectorization of technical drawings. In Computer Vision – ECCV 2020, pages 582–598, 2020 >

[8] Google Fonts. Noto Emoji. https://github.com/googlefonts/noto-emoji, 2026 >

[9] Qijia He, Xunmei Liu, Hammaad Memon, Ziang Li, Zixian Ma, Jaemin Cho, Jason Ren, Daniel S. Weld, and Ranjay Krishna. Vfig: Vectorizing complex figures in svg with vision-language models. arXiv preprint arXiv:2603.24575, 2026 >

[10] Or Hirschorn, Amir Jevnisek, and Shai Avidan. Optimize & reduce: A top-down approach for image vectorization. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 2148–2156, 2024 >

[11] Juncheng Hu, Ziteng Xue, Guotao Liang, Anran Qi, Buyu Li, Sheng Wang, Dong Xu, and Qian Yu. Amodalsvg: Amodal image vectorization via semantic layer peeling. arXiv preprint arXiv:2604.10940, 2026 >

[12] Teng Hu, Ran Yi, Baihong Qian, Jiangning Zhang, Paul L. Rosin, and Yu-Kun Lai. Supersvg: Superpixelbased scalable vector graphics synthesis. arXiv preprint arXiv:2406.09794, 2024 >

[13] Ajay Jain, Amber Xie, and Pieter Abbeel. Vectorfusion: Text-to-svg by abstracting pixel-based diffusion models.

Improvements for AI systems

  1. Bold header: Sparse Anchor Field Prediction for Structural Initialization. The system can model anchor placement before constructing the SVG path by predicting an image-conditioned sparse anchor field, which serves as a learned structural intermediate that guides stable anchor placement before being parsed into an ordered Bézier path.

  2. Bold header: Rendering-Guided Iterative Correction Loop. AnchorFlow introduces a mechanism where rendering errors are fed back to update the anchor field and re-resolve the path when needed, transforming reconstruction from a one-shot prediction into an iterative correction process while preserving an explicit editable structure.

  3. Bold header: Robustness Against Imperfect Evidence. The system can achieve robust editable reconstruction from imperfect path evidence by showing that the learned anchor field supports stable anchor placement even when local path evidence is noisy or boundary perturbed, avoiding overfitting to local artifacts.

  4. Bold header: Parameter-Efficient Editable Output Generation. AnchorFlow produces compact editable SVGs with competitive raster fidelity by focusing on structural anchors rather than dense curve fitting, as evidenced by achieving fewer editable parameters on all three datasets compared to baselines like AdaVec [31].

  5. Bold header: Task-Based Editability Optimization. The system can be optimized for user tasks because the evaluation shows AnchorFlow is preferred over competitors for editability, such as being preferred over AutoTrace [1], LIVE [17], and AdaVec [31] in the editing task.

Sources

Related papers