OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents
Andrea Caciolai, Pere-Lluís Huguet Cabot, Chierh Cheng, Albert Ventayol-Boada, Gabriel Mejia Gonzalez, Christophe Ropers, Lucas Bandarkar, Sebastian Ruder, Darlene Sakakihara, Elliot Yun, Pierre Andrews, Grégoire Mialon, Romain Froger, Marta R. Costa-jussà
cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
Code: https://github.com/facebookresearch/meta-agents-research-environments
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English.
Terminology
Abstract
Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question. We introduce OmnilingualGAIA2, a machine-translated expansion (with partial human- expert validation) of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier. Evaluating seven frontier and open-weight agents, we find a universal cross-lingual gap of 8.8-18.4 pass@3 points that is agent-asymmetric in magnitude, concentrates on tool-orchestration rather than quantitative reasoning, and does not close with model scale. A stratified error attribution decomposes the gap as predominantly model-driven (55%), with a bounded translation-contamination floor of only 6.4% of scenario-language pairs. Human-expert linguistic analysis further identifies morphological cue loss and amplified ambiguity as the primary failure mechanisms in non-Latin-script languages. Our results argue that multilingual agentic evaluation must become a standard part of the reporting protocol for globally deployed agents.
Sources
- Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts
- TelcoAgent-Bench: A Multilingual Benchmark for Telecom AI Agents
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Challenges and Recommendations for LLM-as-a-Judge in Multilingual Settings and for Low-Resource Languages
- Gemini: A Family of Highly Capable Multimodal Models
- Gemma: Open Models Based on Gemini Research and Technology
- GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation
- PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
- Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- SEATauBench: Progressively Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages
- Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
- Omnilingual MT: Machine Translation for 1,600 Languages
- gpt-oss-120b & gpt-oss-20b Model Card
- LLM Evaluators Recognize and Favor Their Own Generations
- Qwen3 Technical Report
- Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
- The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering