PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows
summary
The gist
The paper introduces PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows.
In short
The episode discusses 'PolyWorkBench,' a benchmark for LLM agents that tests cross-lingual, long-horizon workflows across five workplace domains. Hosts analyze how current models struggle with tasks requiring multiple languages and multi-step planning. They conclude that agent evaluation must consider language interaction, not just single metrics.
Key concepts
- Cross-Lingual Long-Horizon Workflows
- These are complex tasks where an LLM agent must perform a multi-step job while juggling multiple source and target languages. The benchmark tests this combination, unlike previous benchmarks that focused on only one aspect.
- Cross-lingual Trajectory Coupling
- A concept describing how languages become interwoven throughout an agent's decision-making process during a long task. Errors can propagate across different languages as the workflow progresses, which is a key failure mode.
- Grade, Pytest, and LLM-as-Judge
- These are three distinct evaluation metrics used in PolyWorkBench. Grade checks for structural completeness (partial credit); Pytest runs automated tests on output; and the Judge assesses overall writing quality and coherence.
Terminology used across episodes
This episode discusses
- PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents · Paper Radio
- WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
- RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
- Measuring Massive Multitask Language Understanding
- Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
- ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
- Language Models are Multilingual Chain-of-Thought Reasoners
- CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
- OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
- OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
- tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
The paper
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents · Read on arXiv
Hongliang Li, Yijin Liu, Zhiwei Zhang, Zihe Liu, Xinyue Lou, Jinan Xu, Fandong Meng, Kaiyu Huang
Beijing Jiaotong University · Weixin AI, Tencent Inc
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows".
Jane: The paper was written by Hongliang Li, Yijin Liu, Zhiwei Zhang, Zihe Liu, Xinyue Lou et al. from Beijing Jiaotong University and Weixin AI, Tencent Inc.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, welcome back to the show, everybody. Today we're digging into a fresh arXiv paper that has both of us genuinely excited. It's called "PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents."
Jane: And Tom, I gotta say, just that title alone tells you a lot. We've got "multilingual" and "long-horizon" and "agents" all in one phrase, and that's a combination we don't see very often.
Tom: Right, because usually benchmarks pick one thing. You either test how well a model handles multiple languages, or you test how well an agent handles a long, complicated task. This paper says, why not both at the same time?
Jane: Exactly. And that's the part that got me hooked. They built sixty-seven tasks across five different workplace domains—commerce, knowledge work, legal, localization, and manufacturing—and every single task forces the agent to juggle multiple languages while doing a long, multi-step job.
Tom: So instead of just translating a sentence or answering a trivia question, the agent has to read a Japanese receipt, do some math, and produce an English spreadsheet. That's a whole different ballgame.
Jane: It really is. And the authors make a point that I think is so important: real workplaces don't happen in one language. You've got instructions in one language, source documents in another, and the final report needs to be in a third.
Tom: And that's the gap they're trying to fill. Most agent benchmarks assume everything happens in English, and most multilingual benchmarks don't involve any tool use or planning. This paper says, let's put those together and see what breaks.
Jane: And spoiler alert, a lot breaks. But we'll get into the numbers in a minute. For now, I just love that they're asking the question nobody else is asking.
Tom: Yeah, and it's a question that's going to matter more and more as these agents actually get deployed in real companies. Speaking of which, I want to bring in Lu, our senior researcher, because I know you've got thoughts on this.
Lu: Oh, I absolutely do, Tom. This paper is essentially saying that language isn't just an input-output property. It's woven through the entire decision-making process. When an agent reasons in one language but has to produce output in another, that's a form of cross-lingual trajectory coupling, and errors can propagate in ways we haven't studied.
Jane: Cross-lingual trajectory coupling, I love that phrase. It sounds fancy, but it just means the languages get tangled up with each other as the task goes along.
Lu: Precisely. And that's why the benchmark is designed the way it is. eighty-eight percent of the tasks involve three or more languages across the instruction, the source materials, and the expected output. So it's not just a translation step bolted onto the end.
Tom: Okay, so we've got a benchmark that's testing something new. But the big question for me is, how do the current models actually do on it? That's the part I can't wait to dig into.
Jane: Same here. But before we jump into the results, let's just sit with the design for a second. The fact that they built this from scratch, hand-authored tasks, no machine translation for the source files, that's a lot of care going into making sure the benchmark actually tests what it claims to test.
Lu: And that's exactly why the results are going to be so telling. When a model fails on this benchmark, it's not failing because of a trick question. It's failing because the real world is messy and multilingual, and we haven't been training for that.
Tom: Alright, I'm sold on the premise. Next segment, we're going to look at what happens when you actually run these models through the wringer. Stay with us.
Summary: Tom: Welcome back. So we've established that "PolyWorkBench" is testing something real and important. Now let's talk about what actually happened when they ran the models.
Jane: And Tom, the headline is pretty stark. The best model in the whole evaluation, Claude Opus four point eight running on the ClaudeCode harness, scored zero point nine two one Pass@one. That sounds decent, but it means even the best system out there is losing almost eight points on average across these tasks.
Tom: And that's the top. Most models scored well below that. Only three entries out of eighteen even broke zero point seven nine. The rest were all under zero point seven seven, and some dipped way down into the zero point five range.
Lu: But here's what really caught my eye, Tom. The harness matters almost as much as the model. Claude Opus four point eight scored zero point nine two one on ClaudeCode but dropped to zero point seven one two on OpenClaw. That's a twenty-one-point swing just from changing the scaffolding around the same model.
Jane: That's wild. It's like having the same driver but putting them in a completely different car. The engine's the same, but the steering and the brakes are different, and that changes everything.
Meng: And from an engineering standpoint, that's a huge red flag for anyone trying to deploy these systems. If your choice of agent framework can move the score by twenty points, then you can't just pick a model and call it done. You have to tune the whole pipeline.
Tom: Exactly, Meng. And the paper makes that point really clearly. They publish the full model times harness matrix instead of collapsing everything into one number, because one number would be misleading.
Jane: Now, the other big finding is about which domains are hardest. And this is where it gets really interesting, because it's not uniform. Commerce is the killer. Models that score zero point eight five to zero point nine zero on Knowledge, Legal, and Manufacturing often collapse to zero point five zero or zero point six three on Commerce.
Lu: And that's not random, Tom. Commerce tasks in this benchmark involve numerical reconciliation, currency conversion, strict spreadsheet schemas. One arithmetic mistake voids the entire deliverable. There's no partial credit for getting close.
Meng: So it's a compounding error problem. You make one small mistake early, and it poisons everything downstream. That's exactly the kind of failure that's most dangerous in real business use.
Jane: And then there's the language dimension. The gap between the best and worst language for a single model can be over thirty Grade points. That's bigger than the difference between adjacent rows on the leaderboard.
Tom: So language isn't just a small penalty on top of everything else. It's a whole separate axis of difficulty that can completely change how well a model performs.
Lu: And the failure modes are different, too. Some are pure comprehension errors, where the model just misreads something in the source language. But others are coordination errors, where the model understands everything but fails to keep the source and target languages aligned across multiple steps.
Jane: That coordination error is the really new finding. It's not something you'd see in a static multilingual benchmark, because there's no trajectory to drift. It only shows up when you have a long task with multiple steps.
Tom: So to summarize where we are: the benchmark is hard, the harness matters, Commerce is brutal, and language variation creates its own failure modes. Next up, we're going to talk about how they actually evaluate these tasks, because that's a whole other layer of cleverness.
Improvements: Tom: Alright, so we know the models struggle. But how do you even grade something like "produce a market analysis report in French based on Korean source data"? That's not a multiple-choice question.
Jane: And that's where the evaluation framework in "PolyWorkBench" gets really clever, Tom. They don't rely on one single metric. They use three different ones, and each one catches something the others miss.
Meng: Yeah, and I appreciate that, because as someone who actually builds these systems, I need to know why a model failed. Is it structurally wrong? Is it functionally broken? Or is it just poorly written?
Jane: Exactly. So first, there's Grade, which is task-specific structural scoring. It checks whether the output has all the required components, and it gives partial credit. So if you get the format right but miss a number, you don't get zero.
Tom: Then there's Pytest, which is the executable verification. This is where they actually run automated tests on the output. Does the spreadsheet have the right schema? Is the arithmetic correct? Does the JSON parse?
Lu: And those two are strongly aligned, which makes sense. Grade gives partial credit on the same structural elements that Pytest verifies deterministically. They're measuring the same thing, just at different granularities.
Meng: But here's the thing that surprised me. The third metric, the LLM-as-Judge, barely correlates with the other two. The correlation with Grade is only zero point one eight, and if you only look at tasks where Grade is above zero point five, the correlation drops to negative zero point zero four.
Jane: That sounds bad at first, but the paper explains why it's actually by design. The Judge is measuring something completely different. It's looking at coherence, fluency, faithfulness to the instruction, overall quality of the writing.
Tom: So you can have a task that's structurally perfect, passes every automated test, but reads like it was written by a robot with no sense of the language. And the Judge catches that.
Lu: And the paper gives concrete examples. Tasks like crosslingual fact-checking, patent prior art analysis, policy briefs. These are long-form generative outputs where structural correctness is decoupled from writing quality. The deterministic checks sign off, but a fluent reader would find it deficient.
Meng: So the Judge is basically a quality gate that catches the stuff that rule-based systems can't see. That's actually a really smart division of labor.
Jane: And there's a reverse direction too. Some tasks have low Grade but high Judge scores. Those are mostly Commerce tasks where the model wrote a beautiful narrative explanation but failed the actual numerical check.
Tom: So the model can talk a good game but not deliver the goods. That's a really important failure mode to catch.
Lu: And that's why the paper argues you need all three metrics together. Grade and Pytest verify the artifact, Judge verifies the communication. Neither alone captures both.
Jane: Now, there's one more thing in the evaluation that I think is really smart, and it's about how they report the final score. They use Pass@one as the main metric, but they also report Pass@three for models that were run multiple times.
Tom: And that Pass@three number tells you something important. For the top model, Claude Opus four point eight, the gap between Pass@one and Pass@three is only zero point zero zero seven. It's already saturated. Running it three times doesn't help because it's already succeeding almost every time.
Meng: But for mid-tier models, the gap is huge. Qwen3 point 6-27B on Hermes gains zero point two zero six points from three samples. That means a lot of its failures are just stochastic noise, not fundamental capability gaps.
Lu: And that's a really practical insight. If you're deploying a mid-tier model, you can get significantly better results just by running it multiple times and picking the best output. That's a cheap win that doesn't require a better model.
Jane: So the evaluation framework isn't just about ranking models. It's about understanding why they fail and what you can do about it. That's the kind of insight that actually helps people build better systems.
Tom: Alright, we've covered the design, the results, and the evaluation. For our final segment, let's zoom out and talk about what this all means for the future.
Conclusion: Tom: So we've spent this whole episode on "PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents," and I think we should wrap up by talking about why this matters beyond just another benchmark.
Jane: Yeah, because benchmarks are only useful if they point us toward real improvements. And I think this one does, because it exposes a gap that nobody else was measuring.
Lu: And that gap is the interaction between language and execution. We've known that models struggle with low-resource languages. We've known they struggle with long tasks. But this paper shows that the combination is worse than the sum of its parts.
Meng: And that's the part that keeps me up at night, honestly. If you're building an agent for a real company, you're going to have documents in multiple languages, and you're going to have multi-step workflows. This benchmark shows that current models are not reliable for that.
Tom: But it's not all doom and gloom, right? The paper also shows that some of the failures are fixable. The Pass@three results suggest that multi-sample verification can recover a lot of the lost performance for mid-tier models.
Jane: And that's a practical lever. You don't need a brand new model architecture to get better results. You just need to run the existing model a few times and pick the best output.
Lu: But the deeper implication is that we need to think about training differently. If language variation is embedded throughout the execution trajectory, then we need to train models on multilingual agentic tasks, not just multilingual text or monolingual agent tasks.
Meng: And we need better evaluation too. The fact that the Judge metric barely correlates with the deterministic metrics tells us that we're missing a whole dimension of quality. We need to figure out how to measure that reliably.
Tom: So what's the takeaway for our listeners? I think it's that the era of monolingual agent benchmarks is over. If you're building or evaluating agents, you need to think about language from the start, not as an afterthought.
Jane: And the paper's authors are planning to keep expanding it. More languages, more domains, more evolving workflows. So this is going to be a living benchmark that tracks progress over time.
Lu: And that's exactly what we need. A moving target that keeps pace with the technology, so we can actually see whether we're making progress on the problems that matter.
Tom: Alright, I think we've given "PolyWorkBench" a proper send-off. It's a benchmark that asks the right questions, even if the answers are uncomfortable for current models.
Jane: And uncomfortable answers are how we learn. Thanks for joining us, everybody. We'll see you next time with a new paper to dig into.
Tom: Take care, folks. Keep reading those arXiv abstracts.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language