Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture

arXiv:2608.28641 · cs.CL, cs.AI · Submitted 2026-08-11 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-11

Code: https://github.com/lilt/terminal-bench-lilt

License: http://creativecommons.org/licenses/by/4.0/

The gist: Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment.

Terminology

Abstract

Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 authentic coding tasks in ten languages: Arabic, Czech, German, Spanish, Hindi, Japanese, Korean, Serbian, Turkish, and Chinese. Each task targets issues specific to non-English software development that have no direct English equivalent, e.g., internationalization, encoding, text normalization, and cultural conventions. All tasks are authored by native-speaker programmers and validated through a multi-stage quality control pipeline. Evaluation of six frontier models reveals that even the strongest model reaches only 63.1% pass rate, with many tasks unsolved by any model. Performance varies substantially by language and does not track general coding benchmark rankings, highlighting that multilingual coding competence is a distinct and underexplored capability axis. Sample tasks are available at https://github.com/lilt/terminal-bench-lilt

Related papers