PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents

arXiv:2607.06008 · cs.AI, cs.CL · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows".

Jane: The paper was written by Hongliang Li, Yijin Liu, Zhiwei Zhang, Zihe Liu, Xinyue Lou et al. from Beijing Jiaotong University and Weixin AI, Tencent Inc.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everybody. Today we're digging into a fresh arXiv paper that has both of us genuinely excited. It's called "PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents."

Jane: And Tom, I gotta say, just that title alone tells you a lot. We've got "multilingual" and "long-horizon" and "agents" all in one phrase, and that's a combination we don't see very often.

Tom: Right, because usually benchmarks pick one thing. You either test how well a model handles multiple languages, or you test how well an agent handles a long, complicated task. This paper says, why not both at the same time?

Jane: Exactly. And that's the part that got me hooked. They built sixty-seven tasks across five different workplace domains—commerce, knowledge work, legal, localization, and manufacturing—and every single task forces the agent to juggle multiple languages while doing a long, multi-step job.

Tom: So instead of just translating a sentence or answering a trivia question, the agent has to read a Japanese receipt, do some math, and produce an English spreadsheet. That's a whole different ballgame.

Jane: It really is. And the authors make a point that I think is so important: real workplaces don't happen in one language. You've got instructions in one language, source documents in another, and the final report needs to be in a third.

Tom: And that's the gap they're trying to fill. Most agent benchmarks assume everything happens in English, and most multilingual benchmarks don't involve any tool use or planning. This paper says, let's put those together and see what breaks.

Jane: And spoiler alert, a lot breaks. But we'll get into the numbers in a minute. For now, I just love that they're asking the question nobody else is asking.

Tom: Yeah, and it's a question that's going to matter more and more as these agents actually get deployed in real companies. Speaking of which, I want to bring in Lu, our senior researcher, because I know you've got thoughts on this.

Lu: Oh, I absolutely do, Tom. This paper is essentially saying that language isn't just an input-output property. It's woven through the entire decision-making process. When an agent reasons in one language but has to produce output in another, that's a form of cross-lingual trajectory coupling, and errors can propagate in ways we haven't studied.

Jane: Cross-lingual trajectory coupling, I love that phrase. It sounds fancy, but it just means the languages get tangled up with each other as the task goes along.

Lu: Precisely. And that's why the benchmark is designed the way it is. eighty-eight percent of the tasks involve three or more languages across the instruction, the source materials, and the expected output. So it's not just a translation step bolted onto the end.

Tom: Okay, so we've got a benchmark that's testing something new. But the big question for me is, how do the current models actually do on it? That's the part I can't wait to dig into.

Jane: Same here. But before we jump into the results, let's just sit with the design for a second. The fact that they built this from scratch, hand-authored tasks, no machine translation for the source files, that's a lot of care going into making sure the benchmark actually tests what it claims to test.

Lu: And that's exactly why the results are going to be so telling. When a model fails on this benchmark, it's not failing because of a trick question. It's failing because the real world is messy and multilingual, and we haven't been training for that.

Tom: Alright, I'm sold on the premise. Next segment, we're going to look at what happens when you actually run these models through the wringer. Stay with us.

Summary: Tom: Welcome back. So we've established that "PolyWorkBench" is testing something real and important. Now let's talk about what actually happened when they ran the models.

Jane: And Tom, the headline is pretty stark. The best model in the whole evaluation, Claude Opus four point eight running on the ClaudeCode harness, scored zero point nine two one Pass@one. That sounds decent, but it means even the best system out there is losing almost eight points on average across these tasks.

Tom: And that's the top. Most models scored well below that. Only three entries out of eighteen even broke zero point seven nine. The rest were all under zero point seven seven, and some dipped way down into the zero point five range.

Lu: But here's what really caught my eye, Tom. The harness matters almost as much as the model. Claude Opus four point eight scored zero point nine two one on ClaudeCode but dropped to zero point seven one two on OpenClaw. That's a twenty-one-point swing just from changing the scaffolding around the same model.

Jane: That's wild. It's like having the same driver but putting them in a completely different car. The engine's the same, but the steering and the brakes are different, and that changes everything.

Meng: And from an engineering standpoint, that's a huge red flag for anyone trying to deploy these systems. If your choice of agent framework can move the score by twenty points, then you can't just pick a model and call it done. You have to tune the whole pipeline.

Tom: Exactly, Meng. And the paper makes that point really clearly. They publish the full model times harness matrix instead of collapsing everything into one number, because one number would be misleading.

Jane: Now, the other big finding is about which domains are hardest. And this is where it gets really interesting, because it's not uniform. Commerce is the killer. Models that score zero point eight five to zero point nine zero on Knowledge, Legal, and Manufacturing often collapse to zero point five zero or zero point six three on Commerce.

Lu: And that's not random, Tom. Commerce tasks in this benchmark involve numerical reconciliation, currency conversion, strict spreadsheet schemas. One arithmetic mistake voids the entire deliverable. There's no partial credit for getting close.

Meng: So it's a compounding error problem. You make one small mistake early, and it poisons everything downstream. That's exactly the kind of failure that's most dangerous in real business use.

Jane: And then there's the language dimension. The gap between the best and worst language for a single model can be over thirty Grade points. That's bigger than the difference between adjacent rows on the leaderboard.

Tom: So language isn't just a small penalty on top of everything else. It's a whole separate axis of difficulty that can completely change how well a model performs.

Lu: And the failure modes are different, too. Some are pure comprehension errors, where the model just misreads something in the source language. But others are coordination errors, where the model understands everything but fails to keep the source and target languages aligned across multiple steps.

Jane: That coordination error is the really new finding. It's not something you'd see in a static multilingual benchmark, because there's no trajectory to drift. It only shows up when you have a long task with multiple steps.

Tom: So to summarize where we are: the benchmark is hard, the harness matters, Commerce is brutal, and language variation creates its own failure modes. Next up, we're going to talk about how they actually evaluate these tasks, because that's a whole other layer of cleverness.

Improvements: Tom: Alright, so we know the models struggle. But how do you even grade something like "produce a market analysis report in French based on Korean source data"? That's not a multiple-choice question.

Jane: And that's where the evaluation framework in "PolyWorkBench" gets really clever, Tom. They don't rely on one single metric. They use three different ones, and each one catches something the others miss.

Meng: Yeah, and I appreciate that, because as someone who actually builds these systems, I need to know why a model failed. Is it structurally wrong? Is it functionally broken? Or is it just poorly written?

Jane: Exactly. So first, there's Grade, which is task-specific structural scoring. It checks whether the output has all the required components, and it gives partial credit. So if you get the format right but miss a number, you don't get zero.

Tom: Then there's Pytest, which is the executable verification. This is where they actually run automated tests on the output. Does the spreadsheet have the right schema? Is the arithmetic correct? Does the JSON parse?

Lu: And those two are strongly aligned, which makes sense. Grade gives partial credit on the same structural elements that Pytest verifies deterministically. They're measuring the same thing, just at different granularities.

Meng: But here's the thing that surprised me. The third metric, the LLM-as-Judge, barely correlates with the other two. The correlation with Grade is only zero point one eight, and if you only look at tasks where Grade is above zero point five, the correlation drops to negative zero point zero four.

Jane: That sounds bad at first, but the paper explains why it's actually by design. The Judge is measuring something completely different. It's looking at coherence, fluency, faithfulness to the instruction, overall quality of the writing.

Tom: So you can have a task that's structurally perfect, passes every automated test, but reads like it was written by a robot with no sense of the language. And the Judge catches that.

Lu: And the paper gives concrete examples. Tasks like crosslingual fact-checking, patent prior art analysis, policy briefs. These are long-form generative outputs where structural correctness is decoupled from writing quality. The deterministic checks sign off, but a fluent reader would find it deficient.

Meng: So the Judge is basically a quality gate that catches the stuff that rule-based systems can't see. That's actually a really smart division of labor.

Jane: And there's a reverse direction too. Some tasks have low Grade but high Judge scores. Those are mostly Commerce tasks where the model wrote a beautiful narrative explanation but failed the actual numerical check.

Tom: So the model can talk a good game but not deliver the goods. That's a really important failure mode to catch.

Lu: And that's why the paper argues you need all three metrics together. Grade and Pytest verify the artifact, Judge verifies the communication. Neither alone captures both.

Jane: Now, there's one more thing in the evaluation that I think is really smart, and it's about how they report the final score. They use Pass@one as the main metric, but they also report Pass@three for models that were run multiple times.

Tom: And that Pass@three number tells you something important. For the top model, Claude Opus four point eight, the gap between Pass@one and Pass@three is only zero point zero zero seven. It's already saturated. Running it three times doesn't help because it's already succeeding almost every time.

Meng: But for mid-tier models, the gap is huge. Qwen3 point 6-27B on Hermes gains zero point two zero six points from three samples. That means a lot of its failures are just stochastic noise, not fundamental capability gaps.

Lu: And that's a really practical insight. If you're deploying a mid-tier model, you can get significantly better results just by running it multiple times and picking the best output. That's a cheap win that doesn't require a better model.

Jane: So the evaluation framework isn't just about ranking models. It's about understanding why they fail and what you can do about it. That's the kind of insight that actually helps people build better systems.

Tom: Alright, we've covered the design, the results, and the evaluation. For our final segment, let's zoom out and talk about what this all means for the future.

Conclusion: Tom: So we've spent this whole episode on "PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents," and I think we should wrap up by talking about why this matters beyond just another benchmark.

Jane: Yeah, because benchmarks are only useful if they point us toward real improvements. And I think this one does, because it exposes a gap that nobody else was measuring.

Lu: And that gap is the interaction between language and execution. We've known that models struggle with low-resource languages. We've known they struggle with long tasks. But this paper shows that the combination is worse than the sum of its parts.

Meng: And that's the part that keeps me up at night, honestly. If you're building an agent for a real company, you're going to have documents in multiple languages, and you're going to have multi-step workflows. This benchmark shows that current models are not reliable for that.

Tom: But it's not all doom and gloom, right? The paper also shows that some of the failures are fixable. The Pass@three results suggest that multi-sample verification can recover a lot of the lost performance for mid-tier models.

Jane: And that's a practical lever. You don't need a brand new model architecture to get better results. You just need to run the existing model a few times and pick the best output.

Lu: But the deeper implication is that we need to think about training differently. If language variation is embedded throughout the execution trajectory, then we need to train models on multilingual agentic tasks, not just multilingual text or monolingual agent tasks.

Meng: And we need better evaluation too. The fact that the Judge metric barely correlates with the deterministic metrics tells us that we're missing a whole dimension of quality. We need to figure out how to measure that reliably.

Tom: So what's the takeaway for our listeners? I think it's that the era of monolingual agent benchmarks is over. If you're building or evaluating agents, you need to think about language from the start, not as an afterthought.

Jane: And the paper's authors are planning to keep expanding it. More languages, more domains, more evolving workflows. So this is going to be a living benchmark that tracks progress over time.

Lu: And that's exactly what we need. A moving target that keeps pace with the technology, so we can actually see whether we're making progress on the problems that matter.

Tom: Alright, I think we've given "PolyWorkBench" a proper send-off. It's a benchmark that asks the right questions, even if the answers are uncomfortable for current models.

Jane: And uncomfortable answers are how we learn. Thanks for joining us, everybody. We'll see you next time with a new paper to dig into.

Tom: Take care, folks. Keep reading those arXiv abstracts.

Hongliang Li, Yijin Liu, Zhiwei Zhang, Zihe Liu, Xinyue Lou, Jinan Xu, Fandong Meng, Kaiyu Huang

Beijing Jiaotong University · Weixin AI, Tencent Inc

cs.AI, cs.CL

Submitted: 2026-08-17

Updated: 2026-08-18

Comments: 15 Pages, 6 figures

Code: https://github.com/tatsu-lab/alpaca_eval

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 82/100

The gist: The paper introduces PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows.

Key concepts

Cross-Lingual Long-Horizon Workflows
These are complex tasks where an LLM agent must perform a multi-step job while juggling multiple source and target languages. The benchmark tests this combination, unlike previous benchmarks that focused on only one aspect.
Cross-lingual Trajectory Coupling
A concept describing how languages become interwoven throughout an agent's decision-making process during a long task. Errors can propagate across different languages as the workflow progresses, which is a key failure mode.
Grade, Pytest, and LLM-as-Judge
These are three distinct evaluation metrics used in PolyWorkBench. Grade checks for structural completeness (partial credit); Pytest runs automated tests on output; and the Judge assesses overall writing quality and coherence.

Terminology

Summary

The paper introduces PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows. The authors argue that most existing evaluation settings implicitly assume that the entire agentic process operates within a single linguistic regime, whereas real-world applications often involve multilingual inputs and outputs within a unified workflow. They identify a key limitation in prior research: multilinguality and agentic reasoning are treated as orthogonal dimensions — multilingual benchmarks study language variation in static settings without execution dynamics, while agent benchmarks study execution complexity under monolingual assumptions.

Benchmark composition. PolyWorkBench consists of 67 tasks across five domains, including commerce, knowledge work, legal analysis, localization, and manufacturing. The five domains are Commerce (16 tasks), Knowledge (11), Legal (15), Localization (11), and Manufacturing (14). Tasks are partitioned into a baseline pool (29 tasks, 6–8 estimated steps) and a harder stress pool (38 tasks, 8–12 steps), jointly spanning difficulty levels 3–6 on a 1–6 scale, with a mean of 8.5 estimated tool-use steps per task and per-task time budgets of 30–60 minutes.

Language coverage. The corpus covers ten languages: English, Chinese, Japanese, Korean, Vietnamese, Russian, French, Spanish, German, and Arabic. Coverage is deliberately trilingual by construction: instruction, source materials, and expected output are drawn from potentially different languages. Specifically, 59 of 67 tasks (88%) involve three or more distinct languages across these three roles. English appears in 66 tasks, Chinese in 62, Korean in 48, Japanese in 47, Vietnamese in 41, Russian in 40, French in 39, German in 38, Spanish in 33, and Arabic in 1 task.

Data construction. Every task is authored by hand rather than crawled or generated in bulk, following a three-stage curation pipeline: (i) source materials collected by domain-familiar annotator[s] in native languages, with machine translation never used to produce a source-language file; (ii) instruction and reference authoring, where ground-truth facts uncovered while writing the reference... are then back-injected into the source files as verifiable anchors; (iii) evaluation assets including a Pytest suite, a weighted grade function with per-dimension rubric, and an LLM-as-Judge prompt. Tasks are admitted only if "the deterministic evaluators separate correct from incorrect trajectories reliably, every ground-truth fact is discoverable from the provided inputs, and no instruction can be satisfied by string-copying from the source without cross-lingual reasoning."

Evaluation framework. The benchmark evaluates along three complementary axes: structural grading, executable verification, and semantic assessment. Grade measures task completion via task-specific structural rules with partial credit on partially completed solutions. Pytest verifies functional correctness on tasks with executable outputs — checking file formats, numerical correctness, schema validity, API outputs, program behaviour. LLM-as-Judge covers coherence, completeness, faithfulness to the task instruction, and overall output quality. The primary leaderboard metric is Pass@1, the average Grade across all 67 tasks under a single decoding sample, with Pass@3 reported for entries with multiple runs.

Experimental setup. The authors evaluate a diverse collection of proprietary and open-weight LLMs across four standardized agent harnesses: ClaudeCode, OpenClaw, Hermes, and Codex, with a timeout budget (1,800 seconds per task). The 18 evaluated model×harness pairs include Claude Opus 4.8, Claude Opus 4.7, DeepSeek-v4-Flash, Qwen3.6-35B-A3B, Qwen3.6-27B, GPT-5.5, Qwen-Agent-World, and GLM-5.1 across various harnesses.

Main results. The best entry, Claude Opus 4.8 + ClaudeCode, reaches 0.921 Pass@1, and only three entries exceed 0.79; every other model falls below 0.77. The authors find that the choice of harness moves scores by 8–21 Pass@1 points on the same underlying model: Claude Opus 4.8 shifts from 0.921 to 0.712 when swapping ClaudeCode for OpenClaw, and Qwen3.6-27B spans 0.157 Pass@1 across three harnesses.

Per-domain analysis. The authors observe a systematic Commerce dip: strong all-around models... retain 0.85–0.90 on Knowledge, Legal, and Manufacturing but collapse to 0.63 and 0.50 on Commerce. They explain that Commerce tasks in PolyWorkBench mix numerical reconciliation, structured spreadsheet output, cross-currency arithmetic, and strict schema constraints, where a single arithmetic or schema mistake voids the entire deliverable. Five of eight Commerce tasks are failed by ≥ 15 of the 18 evaluated agents. Legal has the widest Pass@1 spread across models (0.951 top vs. 0.711 median). Localization is the only domain in which a non-Claude model obtains the second-best score at every tier. Manufacturing has the highest floor: even the weakest entry... still scores 0.638.

Per-language analysis. Strong models like Opus 4.8/ClaudeCode remain balanced across the ten languages (0.82–0.96), while mid-tier models degrade sharply on Russian, Spanish, and German. The gap between best and worst per-language mean can exceed 30 Grade points – larger than the overall Pass@1 differences between adjacent leaderboard rows. Two failure modes recur: pure comprehension errors, in which the agent misreads figures or entities in the source language and never recovers, and cross-lingual coordination errors, in which comprehension is correct but the agent fails to keep source-language content and target-language output aligned across a multi-step trajectory.

Evaluation consistency. Grade and Pytest are strongly aligned (Pearson r = 0.85). However, Judge is only weakly correlated with either signal (r = 0.18 with Grade, r = 0.13 with Pytest). The Judge distribution is heavily bimodal and saturated: 60.1% of scores are ≥ 0.8, 22.5% are ≤ 0.5. When restricting to Grade ≥ 0.5 (81.6% of the population), the Grade–Judge correlation collapses to r = −0.04. The disagreement is asymmetric: 132 pairs have Grade ≥ 0.8 yet Judge ≤ 0.4, versus 97 the other way, with the former clustering on long-form generative outputs (crosslingual fact-checks, contract-conflict analyses, patent prior-art briefs). The authors conclude that Judge is therefore not a suitable ranking metric on its own, but it is the only component sensitive to the semantic degradations that pass every deterministic check.

Sampling headroom. The Pass@3 − Pass@1 gap is small at the top of the leaderboard – Opus 4.8/ClaudeCode gains only +0.007 — and grows monotonically as base capability drops. GPT-5.5/OpenClaw gains +0.141, Qwen3.6-27B/OpenClaw +0.146, Qwen3.6-35B-A3B/Hermes +0.156, and Qwen3.6-27B/Hermes +0.206. The authors conclude that a substantial fraction of their Pass@1 failures are therefore transient stochastic errors rather than fundamental capability gaps.

Harness sensitivity. For every model that has been run under ≥ 2 harnesses, the spread across harnesses is at least 0.08 Pass@1: Opus 4.8 spans 0.209, Qwen3.6-27B spans 0.157, DeepSeek-v4-Flash spans 0.099. In every case, ClaudeCode is either the best or tied-best harness for the model, but the ordering of the remaining three harnesses is not stable across models. The authors conclude that reporting a model's benchmark score without disclosing the harness is therefore not meaningful.

Conclusion. The authors state that multilingual long-horizon workflows remain challenging for current LLM agents and that multilinguality introduces compounding effects across execution trajectories, affecting not only comprehension but also planning stability and tool-use reliability. They argue that future agent evaluation should jointly consider language variation and procedural execution, rather than treating them as independent dimensions. As future work, they plan to continuously expand the benchmark with additional languages, domains, and evolving real-world workflows.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

Implementation: Add a dedicated alignment layer that tracks source-language entities, figures, and facts throughout the entire execution trajectory, not just at input/output boundaries. This layer maintains a running consistency map between what was read in the source language and what is being generated in the target language at every intermediate step.

What the improved system can do: Prevent cross-lingual coordination errors where comprehension is correct but the final artefact drifts because source-language content and target-language output become misaligned across multi-step workflows. This directly addresses the failure mode identified in Section 4.3 and Appendix A.2.

Implementation: For mid-tier models (Pass@1 below 0.80), automatically run 3 independent decoding samples on tasks with high structural complexity (≥8 estimated steps, strict schema requirements). Use a lightweight verifier to check deterministic properties (numerical correctness, schema validity) across samples, and select the best-scoring output. For top-tier models (Pass@1 above 0.90), skip multi-sampling to save compute, as the paper shows gains are negligible (+0.007).

Implementation: Add a specialized verification checkpoint specifically for Commerce workflows that involve numerical reconciliation, currency conversion, or strict spreadsheet schemas. Before finalizing output, run an arithmetic consistency check that validates all cross-currency calculations, verifies row/column totals, and confirms schema compliance. If any check fails, trigger a targeted re-execution of only the failed sub-step rather than the entire task.

Implementation: Train a small calibration head on top of the agent's final output that predicts per-language reliability based on the model's known degradation patterns (e.g., Russian, Spanish, German are consistently weak for mid-tier models). When the calibration head detects low confidence for a specific language, automatically switch to a stronger model for that step, or increase the number of verification passes.

Implementation: Modify the agent's output generation to explicitly optimize for the three-axis evaluation framework (structural grading, executable verification, semantic quality). This means: (a) always producing machine-verifiable structured artefacts (schemas, formats) even when the task allows free-form output; (b) including explicit numerical checkpoints in intermediate reasoning that can be deterministically verified; (c) generating a brief self-assessment note that flags any known semantic weaknesses.

Implementation: When deploying an agent, automatically select the optimal harness for the specific model based on the published model×harness matrix (Table 1). For Claude Opus 4.8, always use ClaudeCode (0.921 vs 0.712 on OpenClaw). For DeepSeek-v4-Flash, prefer ClaudeCode (0.796) over Hermes (0.698) or OpenClaw (0.708). For Qwen models, prefer ClaudeCode or OpenClaw over Hermes.

The improved AI system can:

  • Maintain semantic consistency across languages throughout multi-step workflows, eliminating cross-lingual drift

  • Self-select between single and multi-sample execution based on predicted reliability

  • Recover from arithmetic/schema errors in Commerce tasks without full re-execution

  • Dynamically escalate low-confidence language tasks to stronger models or additional verification

  • Optimize outputs for the evaluation framework that determines leaderboard ranking

  • Automatically choose the best agent harness for the underlying model

These improvements directly target the specific failure modes quantified in the paper: cross-lingual coordination errors, Commerce task collapse, per-language degradation, harness sensitivity, and the Grade-Judge disagreement gap.

Sources

Related papers