WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective
cs.SE, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-20
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- WebSight: A Vision-First Architecture for Robust Web Agents
- I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications
- Deep Reinforcement Learning for Automated Web GUI Testing
- WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
- Temac: Multi-Agent Collaboration for Automated Web GUI Testing
- Understanding Automated Web GUI Testing: An Empirical Study Across Exploration Strategies and State Abstractions
- WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts
- FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
- Multimodal graph representation learning for website generation based on visual sketch
- FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback
- Web-Bench: A LLM Code Benchmark Based on Web Standards and Frameworks
- Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation
- AWorld: Orchestrating the Training Recipe for Agentic AI
- ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
- WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games
- MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered Assistants
- FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties