Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to 100+ Languages via Back-Translation

arXiv:2510.18019 · cs.CL, cs.AI · Submitted 2025-10-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to 100+ Languages via Back-Translation".

Jane: The paper was written by Asim Mohamed and Martin Gubri from African Institute for Mathematical Sciences and Parameter Lab.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: We're looking at a paper titled 'Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to one hundred plus Languages via Back-Translation'.

Jane: That title really hits you right in the face, doesn't it, Tom? It's such a provocative way to start a research paper.

Tom: Asim Mohamed and Martin Gubri are definitely not playing it safe with that opening.

Lu: I think it's brilliant because it immediately challenges the assumption that our current safety tools work everywhere.

Meng: It sounds like they're suggesting that the watermarks we rely on to identify AI text might be failing in certain parts of the world.

Lalam: That's a huge concern for global digital safety, especially for communities that don't use English as their primary language.

Jane: Exactly, and if the watermarking isn't actually multilingual, then we're leaving a massive gap in how we track synthetic content.

Tom: It makes you wonder if the researchers found a way to fix this or if they're just pointing out a disaster.

Lu: They're doing both, actually, by exposing the flaws in the current systems.

Meng: I'm curious if this is a widespread problem or just something that happens in specific edge cases.

Lalam: It seems to be a structural issue that could affect millions of people if we don't address it.

Jane: We'll see exactly how deep this problem goes in the next part of our chat.

Summary: Tom: We've just been talking about the implications of the title, but now let's look at what's actually happening under the hood of 'Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to one hundred plus Languages via Back-Translation'.

Jane: The authors explain that current methods use something called semantic clustering to try and make watermarks work across languages.

Tom: Right, they group similar words together so the watermark signal stays consistent even if you switch from English to French.

Jane: But the paper shows this breaks down when you move into medium- or low-resource languages.

Lu: It's because of how tokenizers work, which is a fascinatingly messy part of AI.

Meng: So the tokenizer is basically chopping up words into pieces that don't make sense for the watermark?

Jane: That's a great way to put it, Meng. In languages like Tamil or Hebrew, the tokenizer often splits words into tiny sub-character units instead of whole words.

Tom: And since the watermark signal is tied to those semantic clusters of whole words, the signal just disappears into those fragments.

Lu: It's like trying to build a mosaic where half your tiles have been shattered into dust.

Meng: If the signal is that fragile, how easy is it for someone to just translate a text to bypass detection?

Jane: It's incredibly easy, according to their findings.

Lalam: This creates a real cultural divide where certain languages are much harder to protect from misinformation than others.

Tom: They actually showed that the watermark strength drops significantly in these languages because the vocabulary is so fragmented.

Jane: We're going to talk about how they actually solve this next, because their solution is pretty ingenious.

Improvements: Tom: We've seen the problem of fragmented tokens, but now we get to the solution proposed in 'Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to one hundred plus Languages via Back-Translation'.

Jane: They've introduced a new method called STEAM, which stands for Search-based Translation-Enhanced Approach for Multilingual watermarking.

Tom: Instead of relying on those broken clusters, STEAM uses back-translation to find the strongest version of the watermark signal.

Lu: I'm obsessed with how they use Bayesian optimization to do this!

Jane: Can you explain that simply for us, Lu?

Lu: Sure, they don't just guess which language to use; they use a mathematical model to search through one hundred thirty-three different candidate languages to find the one that best recovers the watermark.

Meng: That sounds computationally heavy, though. How many translations are we talking about per piece of text?

Jane: They actually capped it at twenty evaluations per input to keep it efficient.

Meng: That's much more manageable for a real-world application.

Tom: And the results they reported are huge, with average gains of plus zero point two three in AUC and a thirty-seven percent increase in their true positive rate.

Lu: It's a massive leap forward for making watermarking actually work globally.

Lalam: It brings a sense of digital equity to the table by ensuring that low-resource languages get the same level of protection as English.

Jane: It even stays robust even if the attacker uses a different translation service than the defender.

Tom: We're almost at the end of our show, so let's wrap this all up.

Conclusion: Tom: We've covered a lot of ground today with 'Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to one hundred plus Languages via Back-Translation'.

Jane: It's been such an eye-opening discussion about the gaps in our current AI safety tools.

Lu: I'm feeling really optimistic that these kinds of search-based methods will lead to much more resilient AI ecosystems.

Meng: From my side, seeing a method that is non-invasive and doesn't require retraining the whole model is a big win for engineers.

Lalam: I think this work ensures that the benefits of AI safety are shared by all cultures, not just a privileged few.

Tom: Thanks to everyone for joining us, and we'll see you next time for the next paper!

Jane: Goodbye, everyone!

Asim Mohamed, Martin Gubri

African Institute for Mathematical Sciences · Parameter Lab

cs.CL, cs.AI

Submitted: 2025-10-20

Updated: 2026-09-11

Code: https://github.com/asimzz/steam

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 89/100

The gist: This paper examines the efficacy of multilingual LLM watermarking, revealing that current methods are not "truly multilingual" because they fail to remain robust against translation attacks in

Key concepts

Semantic Clustering
This method groups similar words together to help watermark signals stay consistent when switching between languages. However, it often fails in medium- or low-resource languages because the signal relies on whole-word clusters that can be broken apart during the text processing stage.
Tokenizer
A tokenizer is a component that breaks text into smaller pieces. In languages like Tamil or Hebrew, tokenizers often split words into tiny sub-character units rather than whole words. This fragmentation causes the watermark signal to disappear, making it easy to bypass detection through translation.
STEAM
Standing for Search-based Translation-Enhanced Approach for Multilingual watermarking, STEAM uses back-translation and Bayesian optimization to find the strongest watermark signal. It searches through 133 candidate languages to recover the signal, significantly increasing the true positive rate for detecting AI-generated text in diverse languages.

Terminology

Summary

This paper examines the efficacy of multilingual LLM watermarking, revealing that current methods are not truly multilingual because they fail to remain robust against translation attacks in medium- and low-resource languages. This research is critical for ensuring content provenance and preventing the spread of misinformation across the world's linguistic diversity, particularly in communities where moderation tools are less effective.

The Limitations of Semantic Clustering

The authors demonstrate that existing multilingual watermarking methods rely heavily on semantic clustering, which groups semantically equivalent tokens into clusters. While this works for high-resource languages, it fails in medium- and low-resource settings due to the uneven coverage of full-word tokens in tokenizers. Because BPE-based tokenizers favor high-frequency data, low-resource languages are often fragmented into subword units that lack semantic meaning.

Consequently, the authors identify a clear positive correlation between the number of full-word tokens in a tokenizer vocabulary and watermark robustness. As coverage decreases, the watermark strength diminishes significantly. This reveals a fundamental limitation: semantic clustering cannot scale effectively beyond high-resource languages because it cannot resolve the issue of subword fragmentation in underrepresented languages.

The STEAM Framework

To overcome these structural limitations, the authors propose STEAM (Search-based Translation-Enhanced Approach for Multilingual watermarking). STEAM is a detection-time method that is non-invasive, model-agnostic, and compatible with any existing watermarking technique or tokenizer. Instead of relying on fixed semantic clusters, STEAM uses Bayesian optimization to search through 133 candidate languages to find the specific back-translation that best recovers the watermark strength.

This approach allows for fairer watermarking across the diversity of languages by making detection resilient to the tokenization biases found in underrepresented languages. Because it is a detection-time intervention, it does not require changing how models generate text, making it easily extendable to new languages.

Technical Implementation

The STEAM pipeline operates through a sophisticated search mechanism designed to be tractable at scale. The method follows these core steps:

  • Each candidate language is represented by a 131-dimensional feature vector containing syntactic and phonological properties from the URIEL knowledge base.

  • A Bayesian optimization loop uses a Gaussian process surrogate to model the relationship between linguistic features and observed z-scores.

  • The system selects the next back-translation language by maximising the expected improvement over all unevaluated candidates, with evaluations capped at 20 per input.

To address language token bias—where sub-character tokens in low-resource languages can artificially inflate or deflate z-scores—STEAM replaces the standard green token fraction with a language-specific gamma derived from a calibration set of human-written texts.

Performance and Robustness

Experimental results indicate that STEAM provides large and consistent performance gains, achieving average improvements of +0.23 AUC and +37%p TPR@1%. The method proves highly stable across diverse adversarial scenarios:

  • It is robust to translator mismatch, maintaining high AUC even when the attacker and defender use different translation services.

  • It withstands adaptive adversarial evaluation, such as multi-step translation attacks involving pivot languages.

  • It maintains high detection strength across various text lengths, from short to long documents, and remains effective even when the best language is not in its initial candidate pool.

Improvements for AI systems

1. Implementation of Bayesian Optimization-driven Back-Translation (STEAM) in Content Provenance Pipelines

  • Improved AI System Capability: The system can reliably detect LLM-generated synthetic content even after an adversary has attempted to scrub the watermark by translating the text into a medium- or low-resource language. By using a Gaussian Process surrogate to search through 133+ candidate languages using 131-dimensional URIEL linguistic features, the system can identify the optimal back-translation language to recover and validate watermark strength within a highly efficient budget (e.g., 20 evaluations per input).

2. Integration of Language-Specific Null Hypothesis Calibration (gamma) in Detection Algorithms

  • Improved AI System Capability: The system eliminates detection bias and prevents false positives/negatives caused by tokenizer fragmentation. Instead of assuming a uniform green-token fraction across all languages, the system uses an empirical gamma (calculated from human-written calibration sets) to account for sub-character token concentration in low-resource languages (e.g., Tamil, Bengali, or Hebrew). This ensures watermark detection is statistically valid regardless of how a specific tokenizer splits a language's vocabulary.

3. Deployment of Tokenizer-Agnostic and Model-Agnostic Watermark Recovery

  • Improved AI System Capability: The system provides a non-invasive defense layer that can be retroactively applied to any existing watermarking scheme (like KGW or SIR) and any LLM architecture without requiring retraining or modifications to the generation logits. This allows for a scalable, universal safety layer that maintains high AUC (>0.96) and TPR@1% across diverse linguistic families, even when the attacker uses different translation services than the defender.

Abstract

Multilingual watermarking aims to make large language model (LLM) outputs traceable across languages, yet current methods still fall short. Despite claims of cross-lingual robustness, they are evaluated only on high-resource languages. We show that existing multilingual watermarking methods are not truly multilingual: they fail to remain robust under translation attacks in medium- and low-resource languages. We trace this failure to semantic clustering, which fails when the tokenizer vocabulary contains too few full-word tokens for a given language. To address this, we introduce STEAM, a detection method that uses Bayesian optimisation to search among 126 candidate languages for the back-translation that best recovers the watermark strength. It is compatible with any watermarking method, robust across different tokenizers and languages, non-invasive, and easily extendable to new languages. With average gains of +0.23 AUC and +37%p TPR@1%, STEAM provides a scalable approach toward fairer watermarking across the diversity of languages.

Sources

Related papers