Compositional SVG Generation via VLM-Driven Hierarchical Semantic Parsing

arXiv:2609.14657 · cs.CV, cs.AI · Submitted 2026-09-13 · Read on arXiv

cs.CV, cs.AI

Submitted: 2026-09-13

Updated: 2026-09-13

Comments: 26 pages, Accepted to EMNLP 2026 (Main)

License: http://creativecommons.org/licenses/by/4.0/

The gist: While Vision-Language Models (VLMs) excel at visual reasoning, generating structured, editable Scalable Vector Graphics (SVG) remains a fundamental challenge.

Terminology

Abstract

While Vision-Language Models (VLMs) excel at visual reasoning, generating structured, editable Scalable Vector Graphics (SVG) remains a fundamental challenge. Existing pipelines predominantly yield flat, semantically agnostic collections of paths, where editing a single object requires manually identifying its constituent paths. To address this, we propose a VLM-driven agentic framework for semantic compositional SVG generation. Our pipeline recursively parses visual scenes into semantic and geometric hierarchies via top-down decomposition, visual grounding, and prompt-driven amodal occlusion recovery, ensuring each component is geometrically complete. Furthermore, we introduce the Semantic SVG Benchmark with human-annotated semantic groups and novel sub-component metrics (Semantic Recall/Precision, PERE) to explicitly evaluate structural compositionality and functional editability. Experiments show that our natively predicted structures surpass the upper bounds of existing flat-generation methods in both grouping quality and editability, while maintaining state-of-the-art visual fidelity.

Related papers