WorldBench: Evaluating LLMs on Three.js Voxel World Generation
cs.GR, cs.AI, cs.CL, cs.CV, cs.SE
Submitted: 2026-10-07
Updated: 2026-10-07
Code: https://github.com/KrishBakshi/worldbench
Terminology
Sources
- Text-to-Scene with Large Reasoning Models
- GT23D-Bench: A Comprehensive General Text-to-3D Generation Benchmark
- Teaching Large Language Models to Self-Debug
- SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation
- T$^3$Bench: Benchmarking Current Progress in Text-to-3D Generation
- SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code
- WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis
- LLM Evaluators Recognize and Favor Their Own Generations
- SceneEval: Evaluating Semantic Coherence in Text-Conditioned 3D Indoor Scene Synthesis
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- SceneX: Procedural Controllable Large-scale Scene Generation
Related papers
- SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control
- CADReasoner: Iterative Program Editing for CAD Reverse Engineering
- QuadLink: Autoregressive Quad-Dominant Mesh Generation via Point-Relation Learning
- DrawVideo: Grounded and Faithful Multi-Shot Video Generation from Storyboard Keyframe Sketches
- MotionPersona: Real-Time Locomotion Control across Personas, Bodies, and Styles
- MeshSplatBench: A Unified Benchmark for Triangle- and Mesh-Based Neural Rendering