SEATauBench: Progressively Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages
cs.CL, cs.AI
Submitted: 2026-06-27
Updated: 2026-09-05
Comments: EMNLP Findings 2026
Code: https://github.com/SEACrowd/SEATauBench
License: http://creativecommons.org/licenses/by/4.0/
The gist: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood despite its importance to sovereign AI.
Terminology
Abstract
While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood despite its importance to sovereign AI. To fill this gap, we introduce SEATauBench, the first agent-focused evaluation framework for SEA sovereign AI. It adapts Tau2-Bench to five languages---Mandarin, Vietnamese, Thai, Indonesian, and Filipino---and evaluates agents across progressively localized settings that vary the language of user-agent interaction, tool specifications, and task domains. Across three models, we find that English agent capabilities transfer reasonably well when only the conversation language changes, but quality and robustness degrade sharply as more task contexts are localized, with the largest losses in full domain adaptation. We also highlight the limits of English-only agent assessment for predicting agent capabilities in SEA languages. More broadly, SEATauBench provides a diagnostic benchmark and reusable adaptation pipeline for building reliable multilingual agents for linguistically diverse regions. Data and code can be accessed at github.com/SEACrowd/SEATauBench
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Anthropogenic Regional Adaptation in Multimodal Vision-Language Model
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- Kimi K2.5: Visual Agentic Intelligence
- Qwen3 Technical Report
- $\tau$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
- CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
- Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
- $\tau$-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- OpenAI GPT-5 System Card
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering