SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation
cs.AI, cs.CV
Submitted: 2026-09-03
Updated: 2026-09-03
Project page: https://potpov.github.io/svg-score-webpage
License: http://creativecommons.org/licenses/by/4.0/
The gist: Scalable Vector Graphics (SVG) generation is attracting increasing attention as generative models improve in expressiveness and controllability.
Terminology
Abstract
Scalable Vector Graphics (SVG) generation is attracting increasing attention as generative models improve in expressiveness and controllability. Progress, however, is held back by the lack of domain-specific evaluation protocols: current practice relies on metrics designed for natural images, most notably CLIPScore, which was never trained on vector graphics and aligns only partially with human judgment. We introduce, a human-aligned evaluation framework for text-to-SVG generation. Through controlled caption and image perturbations, we first show that CLIP-based scores barely react to the errors SVG generators actually make, such as wrong colors, counts, and spatial relations, and that off-the-shelf Vision-Language Model (VLM) judges, while more sensitive, respond unevenly across error types and SVG styles. We then introduce a human-annotated dataset for Semantic Alignment, measuring how faithfully a generated SVG reflects its caption. Building on it, we develop two complementary evaluators: CLIP scorers adapted to vector graphics and then aligned to human preferences, for fast large-scale evaluation, and a VLM judge trained with supervised fine-tuning and reward-shaped reinforcement learning, for more expressive and interpretable assessment. Using both, we benchmark major open-source, commercial, and optimization-based SVG generators on an independent caption set.
Sources
- SVGEditBench V2: A Benchmark for Instruction-based SVG Editing
- GPT-4o System Card
- Gemma 3 Technical Report
- Rendering-Aware Reinforcement Learning for Vector Graphics Generation
- VectorGym: A Multi-Task Benchmark for SVG Code Generation, Sketching and Editing
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- OpenAI GPT-5 System Card
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation
- Qwen3 Technical Report
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection