Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

arXiv:2608.11110 · cs.CL · Submitted 2026-08-13 · Read on arXiv

Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram

Microsoft Research India

cs.CL

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Accepted in COLM 26

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: This paper investigates whether tool-using language model agents execute the same action policy for semantically identical tasks across different languages.

Terminology

Summary

This paper investigates whether tool-using language model agents execute the same action policy for semantically identical tasks across different languages. The authors argue that for agents, the route is the product: the sequence of actions determines cost, latency, failure modes, and auditability, so two language versions that agree on every final answer can still differ substantially in price, failure mode, and whether safeguards hold across languages. The paper states: Answer agreement is not behavioural agreement. Given the same task in a different language, current tool-using models take a measurably different route.

The study encompasses:

  • 8 models: Gemma-3-27B, Sarvam-M (24B, Indic-specialised), Qwen3-235B-A22B, Llama-4-Maverick (17B active, 128 experts), GPT-OSS-120B, Gemma-3-4B, Qwen3-8B, and Aya-Expanse-8B

  • 6 parallel benchmarks: FLORES-200, XQuAD, XNLI, Belebele, XCOPA, and a purpose-built synthetic benchmark

  • 41 languages across 17 families and 16 scripts

  • 2,382,875 rollouts across 505 cells

  • 809 language pairs

The naive measurement of cross-lingual trace similarity fails due to five confounds, each large enough to change conclusions:

  1. C1 - No baseline: a model asked the same question twice in one language does not answer it the same way, so there is no reference point to measure against. Same-language self-consistency Iwithin = 0.63–0.80, not ≈1.

  2. C2 - Trace length: Short traces score higher than long ones. This confound flips the sign of two interventions and has R2 = 0.67 across cells.

  3. C3 - Empty traces: Empty traces score perfectly (1.0). This collapses one model's headline invariance from 0.79 to 0.02.

  4. C4 - Ceiling: The gap is capped by each model's own reproducibility. Correlation between self-consistency and measured gap is r = +0.97 at T=0.

  5. C5 - Chance floor: unrelated traces already agree by chance more than half the time. Measured c ≈ 0.56, up to 0.95 on short traces.

The paper defines:

  • Iwithin: Same-language trace agreement between two replicates (cross-seed)

  • Icross: Cross-language trace agreement between two replicates

  • Gap Δ = Iwithin − Icross

  • Normalised policy retention Ĩ = Icross/Iwithin: the share of a model's own reproducibility that survives a change of language

Every comparison is length-matched in both directions, pairs are dropped when either trace is empty, and intervals are task-level bootstraps.

Every correction we apply makes it larger, never smaller. The uncorrected baseline gap is +0.0625; with all corrections at T=0, it rises to +0.2074. The paper states: Sampling noise was masking the language effect, not producing it.

Under greedy decoding (T=0), four very different frontier models converge on nearly the same value:

  • Gemma-3-27B: Ĩ = 0.7276

  • Sarvam-M: Ĩ = 0.7076

  • Qwen3-235B: Ĩ = 0.7331

  • Llama-4-Maverick: Ĩ = 0.7092

each retaining 71–73% of its action policy when the language changes. The relative spread falls from 16.1% at T=0.5 to 3.5% at T=0. Model identity explains only 5.7% of variance across 24 cells, while benchmark explains 26.9%.

Cross-lingual agreement is essentially flat across T ∈ 0, 0.3, 0.5, 0.7, 1.0, while self-consistency falls 21× faster. Cross-lingual policy divergence is invariant to sampling temperature.

Below roughly 10B parameters, the regularity breaks down. The within-family comparison shows Gemma-3-4B and Gemma-3-27B (same recipe, 6.75× scale difference) sit ten points apart. The apparent ordering among smaller models is largely an artifact of a chance floor that we measure by permutation rather than assume. The chance correction reverses which of two smaller models looks better.

Agents route non-English tasks through English:

  • Translate is the most-used tool in every adapted benchmark

  • Reasoning text is ≈99% ASCII even on Devanagari input

  • Removing the translation tool lowers length-matched agreement in proportion to usage

  • Mandating it helps in a pre-registered ordering across four models (5/6, 5/6, 3/6, 2/6 benchmarks, monotone in head-room)

  • The pivot survives a direct instruction to abandon it, at under 1% compliance

A single trace-extraction regex manufactured an apparent multilingual failure in GPT-OSS-120B. The model simply answers in prose that names the tool, writing 'We will use Translate.' instead of emitting the required syntax. Two worked exemplars raised measured accuracy twenty-sixfold (0.0174 → 0.4539) while accuracy on readable outputs barely moved (0.8127 → 0.7418). The intervention made it legible, not smarter. The paper recommends treating any model above 20% parse-failure rate as unranked.

  • Correctness: Cross-lingual accuracy spread reaches 0.66 within a single model and benchmark. Every model has an English advantage; Sarvam-M (Indic-specialised) has the largest at +0.155.

  • Invariance is not a proxy for accuracy: The pooled correlation r = +0.897 is an artifact of two clusters; among adherent models it falls to r = +0.378 and reverses in a quarter of cells.

  • Self-consistency voting: Voting costs 1.6–1.9 points of Ĩ with disjoint intervals—a variance reducer, not a retention improver.

  • Trace length is causal: A three-level manipulation moves Ĩ by 6–7 points with disjoint intervals, further than the entire across-model band.

Improvements for AI systems

Based on the paper’s findings, here are specific improvements to AI systems and the resulting capabilities:

Improvement: Implement a runtime monitor that compares the action trace (tool calls, reasoning steps) of a non-English task against the English-equivalent trace at the same length, using the normalized retention metric (Ĩ = Icross/Iwithin). Flag any task where Ĩ drops below 0.71.

Resulting capability: The system can detect when it silently changes its problem-solving strategy across languages—even when final answers match—and alert operators to potential safety or cost regressions before deployment.

Improvement: Since cross-lingual policy divergence is invariant to temperature (while self-consistency drops 21× faster), use language-pair agreement as a stable, temperature-independent quality signal for agentic tasks. Stop tuning temperature for consistency; instead, validate policy retention across languages at T=0.

Improvement: For non-English tasks, either (a) mandate a translation step with a dedicated tool and verify the reasoning text is >95% ASCII, or (b) explicitly disable translation and measure the resulting Ĩ drop. Do not leave the pivot implicit—the model will do it anyway (at <1% compliance with instructions to abandon it).

Improvement: Treat any model with >20% trace-extraction parse failures as unranked. Before comparing models or languages, require a structured-output parser with ≥80% success; otherwise, report the model as “not evaluable” rather than assigning a score.

Improvement: When benchmarking any agent across languages or interventions, always (a) measure same-language self-consistency (Iwithin) as a baseline, (b) length-match traces in both directions, (c) drop empty traces, and (d) compute chance floor via permutation. Report Ĩ, not raw accuracy or raw trace similarity.

Improvement: For models below 10B parameters, add a dedicated diagnostic that measures Ĩ against a frontier model (e.g., 27B+) on the same benchmark. If the small model’s Ĩ is within 3 points of the frontier’s, treat it as policy-compatible; otherwise, flag it as needing architectural changes (e.g., more capacity for cross-lingual abstraction).

Improvement: Instead of voting over multiple same-language samples (which costs 1.6–1.9 points of Ĩ), vote over samples from different languages for the same task. This preserves policy retention while still reducing variance, and it doubles as a multilingual robustness check.

Abstract

When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: they fix cost and latency, decide how the system fails, and are the only auditable part of its behaviour. We make the action policy the measured object across 8 models, 6 parallel benchmarks and 41 languages (2.38M rollouts). The naive measurement fails: five confounds sit between raw trace similarity and any defensible claim, each able to flip a conclusion. Short traces score higher, empty traces score perfectly, unrelated traces agree by chance over half the time, the gap is capped by each model's reproducibility, and a model asked the same question twice in one language answers differently, leaving no baseline. We remove all five, and every correction makes the effect larger. Divergence proves structural, not sampling noise: it survives greedy decoding in every cell and stays flat as temperature rises, even as models grow less self-consistent. Normalised by their own reproducibility, four very different frontier models converge under greedy decoding, each keeping 71-73% of its action policy across languages, with model identity explaining only 5.7% of the variance. Below roughly 10B parameters it breaks down, and the ordering among smaller models is largely an artifact of a chance floor we measure by permutation rather than assume. Agents route non-English tasks through English; this pivot is causally load-bearing, confirmed by a pre-registered prediction across four models, and models will not abandon it when told to. Finally, a single trace-extraction regex, not the model, manufactured a multilingual failure: two worked examples raise one model's measured accuracy twenty-sixfold while its accuracy on readable outputs barely moves.

Sources

Related papers