Enabling Creative Exploration for Vibe Design Agents
cs.AI
Submitted: 2026-09-14
Updated: 2026-09-25
Comments: 14 pages, 3 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vibe design agents turn natural-language briefs into rendered interfaces and frontend code.
Terminology
Abstract
Vibe design agents turn natural-language briefs into rendered interfaces and frontend code. Yet a useful design agent should do more than produce one valid page: it should help users explore coherent alternatives. Increasing token-level temperature is a blunt solution because it varies aesthetic decisions and syntax-sensitive code at the same time. We instead separate exploration from implementation through an inference architecture that makes design direction an explicit intermediate decision. Inspired by Verbalized Sampling, a pre-pass proposes structured design specifications with typicality scores, an external selector samples one, and the downstream generator realizes the selected specification together with the original request under fixed settings. We apply this approach to UI themes and visual-asset prompts. Across 168 prompts, with 1,255 paired comparisons per temperature for each intervention, theme sampling broadens observed selection coverage and screenshot variation, while LLM-judge preferences vary across interventions, prompt complexity, and viewport. In an online experiment with more than 300,000 tasks, the observed code-export increase remains statistically uncertain, while fewer negative feedback events coexist with more correction interactions and modest operational costs. Together, these findings identify structured design specifications as a practical control point for exploring alternative UI concepts while keeping downstream generation settings fixed.
Sources
- The Human Creativity Benchmark
- Diverse Preference Optimization
- Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
- FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs
- Calibrating Verbalized Probabilities for Large Language Models
- DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
- ComUICoder: Component-based Reusable UI Code Generation for Complex Websites via Semantic Segmentation and Element-wise Feedback
- MAxPrototyper: A Multi-Agent Generation System for Interactive User Interface Prototyping
- Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
- Frontend Diffusion: Exploring Intent-Based User Interfaces through Abstract-to-Detailed Task Transitions
- FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection