WorldBench: Culturally Grounded Benchmark for Multilingual Agents
cs.AI, cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite the growing use of LLM-powered agents to solve multi-step tasks in complex environments, existing benchmarks rarely test state preservation, performance across languages, and application to
Terminology
Abstract
Despite the growing use of LLM-powered agents to solve multi-step tasks in complex environments, existing benchmarks rarely test state preservation, performance across languages, and application to realistic, grounded scenarios. To address these concerns, we present WorldBench: a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions. WorldBench comprises 1,600 tasks across seven languages and eight cultures, filtered and refined through feedback from human annotators with language- and culture-specific expertise. For evaluation, we extend metrics from previous works and introduce Constrained Task Success (CTS), which combines natural language instructions and testbeds to score task completion, minimal modification, and other complementary metrics through deterministic and LLM-as-a-Judge evaluations. Our experiments show that frontier models reach only 49.2% CTS, with all models demonstrating large gaps between correctness and environment preservation. We thereby show that current agents remain brittle in multilingual, agentic scenarios, especially for long-horizon tasks and under state-preservation constraints
Sources
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- Mind2Web: Towards a Generalist Agent for the Web
- MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning
- AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping
- Lean-STaR: Learning to Interleave Thinking and Proving
- AgentBench: Evaluating LLMs as Agents
- AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
- ParseBench: A Document Parsing Benchmark for AI Agents
- DocVQA: A Dataset for VQA on Document Images
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
- Self-Instruct: Aligning Language Models with Self-Generated Instructions
- OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- WizardLM: Empowering large pre-trained language models to follow complex instructions
- InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
- From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection