DesignAsCode: Bridging Structural Editability and Visual Fidelity in Graphic Design Generation
cs.GR, cs.AI, cs.CV, cs.LG, cs.MM
Submitted: 2026-02-06
Updated: 2026-08-29
Project page: https://liuziyuan1109.github.io/design-as-code
License: http://creativecommons.org/licenses/by/4.0/
The gist: Graphic design generation demands a delicate balance between high visual fidelity and fine-grained structural editability.
Terminology
Abstract
Graphic design generation demands a delicate balance between high visual fidelity and fine-grained structural editability. However, existing approaches typically bifurcate into either non-editable raster image synthesis or abstract layout generation devoid of visual content. Recent combinations of these two approaches attempt to bridge this gap but often suffer from rigid composition schemas and unresolvable visual dissonances (e.g., text-background conflicts) due to their inexpressive representation and open-loop nature. To address these challenges, we propose DesignAsCode, a novel framework that reimagines graphic design as a programmatic synthesis task using HTML/CSS. Specifically, we introduce a Plan-Implement-Reflect pipeline, incorporating a Semantic Planner to construct dynamic, variable-depth element hierarchies and a Visual-Aware Reflection mechanism that optimizes the code to rectify rendering artifacts. Extensive experiments demonstrate that DesignAsCode significantly outperforms baselines in both structural validity and aesthetic quality. Furthermore, our code-native representation unlocks advanced capabilities, including automatic layout retargeting, complex document generation (e.g., resumes), and CSS-based animation. Our project page is available at https://liuziyuan1109.github.io/design-as-code/.
Sources
- PSD2Code: Automated Front-End Code Generation from Design Files via Multimodal Large Language Models
- DesignCoder: Hierarchy-Aware and Self-Correcting UI Code Generation with Large Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- OmniPSD: Layered PSD Generation with Diffusion Transformer
- PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography
- GPT-4o System Card
- Decoupled Weight Decay Regularization
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- COLE: A Hierarchical Generation Framework for Multi-Layered and Editable Graphic Design
- Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- LayoutGAN: Generating Graphic Layouts with Wireframe Discriminators
- StarVector: Generating Scalable Vector Graphics Code from Images and Text
- PosterLLaVa: Constructing a Unified Multi-modal Layout Generator with LLM
- Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition
- CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design
- Denoising Diffusion Implicit Models
- LayoutNUWA: Revealing the Hidden Layout Expertise of Large Language Models
- CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation
- BannerAgency: Advertising Banner Design with Multimodal LLM Agents
Related papers
- SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control
- CADReasoner: Iterative Program Editing for CAD Reverse Engineering
- QuadLink: Autoregressive Quad-Dominant Mesh Generation via Point-Relation Learning
- DrawVideo: Grounded and Faithful Multi-Shot Video Generation from Storyboard Keyframe Sketches
- MotionPersona: Real-Time Locomotion Control across Personas, Bodies, and Styles
- MeshSplatBench: A Unified Benchmark for Triangle- and Mesh-Based Neural Rendering