Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics

arXiv:2608.10678 · cs.CL, cs.AI · Submitted 2026-08-11 · Read on arXiv

Tsinghua University · SiliconProspect AI · Nanyang Technological University · Alibaba Group

cs.CL, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-31

Code: https://github.com/qingjiesjtu/SampledBPE

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: This paper presents SAMPLED-BPE, a lightweight token-level auditing pipeline for web-scale Chinese corpora, motivated by observations of Chinese web pollution surfacing in LLMs—including "spam- and

Terminology

Summary

This paper presents SAMPLED-BPE, a lightweight token-level auditing pipeline for web-scale Chinese corpora, motivated by observations of Chinese web pollution surfacing in LLMs—including spam- and pornography-related Chinese tokens in ChatGPT's vocabulary and Chinese gambling content in Codex outputs. The authors propose a method that samples a small subset and trains BPE tokenizer to surface polluted tokens, addressing three challenges: (1) their web-scale size makes full scan costly; (2) prior analyses are often too coarse to expose token-level pollution; (3) Chinese web pollution is implicit and rapidly changing.

The SAMPLED-BPE pipeline consists of four stages:

  1. Streaming sampling: we sample documents in a single sequential pass, which makes the sampling cost nearly linear in the input scan and reduces I/O and avoids full materialization of large corpus.

  2. BPE training & counting: We train a BPE tokenizer on the selected corpus and use the same tokenizer to count token frequencies. BPE is chosen because its learned vocabulary surfaces recurrent lexical patterns that characterize the corpus, high-frequency polluted tokens can be exposed without relying on a fixed keyword list, and BPE is efficient for web-scale auditing.

  3. Category mapping: Tokens are mapped into one of six content categories: Normal Content, Adult Content, Online Gambling, Online Gaming, Online Video, and Anomalous. This uses GLM-4-32B, an open-source Chinese LLM with strong Chinese comprehension ability, as the base category classifier, fine-tuned on expert annotations of GPT Chinese vocabularies in Zhang et al. (2025), achieving 97.32% classification accuracy on the test split.

  4. Corpus profiling: We combine token frequencies with category assignments to build a corpus profile, recording each token's ratio, predicted category, Internet evidence, and classification rationale.

Accuracy preservation: "At the 0.25% sampling rate, token coverage remains: the sampled corpus recovers 76.83% of Full tokens. More importantly, the ratios of the covered tokens closely match Full, with Spearman reaching 86.88% and Pearson reaching 99.93%. The weighted relative error remains around 5% at the lowest sampling rate."

Cost reduction: At the 0.25% sampling rate, the auditing pipeline achieves a 148.4× runtime speedup. The memory reduction is also substantial, with peak RSS reduced by 35.8×. The authors state The time cost of auditing 1TB corpus can be reduced from months to hours.

The audit of 11 open Chinese corpora reveals widespread but uneven pollution. Key findings from Table 1:

  • OSCAR has the highest total pollution at 83.38%, dominated by Adult Content at 76.10%

  • mC4 is heavily polluted at 25.41%, with Online Gambling alone reaching 23.17%

  • HPLT contains 4.95% total pollution with both Adult Content and Online Gambling at non-trivial levels

  • CulturaX at 3.36%, CWT at 2.35%, ROOTS at 2.02%

  • Cleaner corpora: WanJuan (0.75%), MAPCC (0.69%), SkyPile (0.61%), CCI3 (0.55%), WuDao (0.50%)

The paper notes: "Broad multilingual web pipelines such as OSCAR, mC4, HPLT, and CulturaX have the highest pollution ratios... In contrast, corpora such as WanJuan, MAPCC, SkyPile, CCI3, and WuDao, which emphasize Chinese-specific collection, trusted sources, or quality filtering, are much cleaner."

Token sharing analysis: Across the 11 corpora, we observe 106,671 polluted tokens. Among them, 102,816 (96.386%) appear in only one corpus. Only 1,542 (1.446%) appear in at least three. The paper finds a small shared core recurring across multiple datasets—252 polluted tokens appear in at least 8 of 11 corpora, spanning all five polluted categories. Examples include 老司机 (old driver) as a persistent euphemism in adult-resource pages and 威尼斯人 (Venetian) as repeatedly reused as one of the world's largest casino resorts.

Auditing six Chinese Common Crawl snapshots from 2021 to 2026 at 1% sampling rate reveals:

From Table 2, total pollution ranges from 34.83% (2024) to 79.52% (2026). Adult Content is both high and volatile, dropping from 55.86% in 2023 to 28.01% in 2024 before rising to 68.72% in 2026. Online Gambling decreases from 5.31% in 2021 to 0.21% in 2026, a trend that is consistent with intensified enforcement against cross-border gambling activities targeting Chinese users. Online Video gradually increases from 3.93% to 8.39%.

The paper states: Chinese web content is highly polluted, and its pollution profile shifts over time. Token evolution analysis shows tokens in Normal Content, Online Gaming, and Online Video are replaced less frequently over time, while Adult Content, Online Gambling, and Anomalous show lower adjacent-year overlap and much weaker five-year persistence, suggesting Adult Content and Online Gambling sites continually reappear on the Chinese web and change their surface forms even after being repeatedly banned or blocked.

The authors release a hierarchical Chinese web token dataset with 630,684 token records, each with web context, category, and explanation fields, organized as 92,972 trees. The hierarchy serves two purposes:

  1. Tracing pollution families: when all tokens in a tree are from the same pollution family, the root identifies a minimal recurring subtoken, enabling summarization of a set of surface variants with a shorter representative token.

  2. Revealing compositional pollution: A subtoken may be normal in isolation, but become polluted when combined with another normal token. Example: 菲律宾 (Philippines) and 申博 (Apply for a PhD) appear as normal tokens separately, while 菲律宾申博 (An online gambling brand) is gambling-related polluted content.

The paper concludes: "Motivated by Chinese web pollution surfacing in ChatGPT's vocabularies and Codex's outputs, this work presents SAMPLED-BPE, a lightweight token-level auditing pipeline that trains BPE tokenizers on sampled corpora to estimate Chinese corpora pollution profiles. Experiments show that SAMPLED-BPE substantially reduces runtime and memory while preserving usable estimates. We audit 11 open Chinese corpora and 6 Chinese Common Crawl snapshots. Results show that pollution is widespread but uneven across open Chinese corpora, and upstream Chinese web content is highly polluted and temporally shifting."

The authors acknowledge: This work focuses on polluted tokens in Chinese web-scale corpora. We do not investigate polluted tokens in other languages, nor do we claim that our auditing pipeline transfers directly beyond Chinese. They also note readability challenges: This paper necessarily includes many Chinese tokens, some of which are offensive, euphemistic, abbreviated, or context-dependent, and translations are readability aids rather than complete semantic equivalents.

Improvements for AI systems

Based on this paper, I can make the following specific improvements to AI systems:

  • Improvement: Integrate SAMPLED-BPE as a pre-training quality gate that automatically samples, tokenizes, and profiles any new web-scale corpus before it enters the training pipeline.

  • What the improved AI can do: Automatically flag and quantify pollution ratios (adult, gambling, etc.) in any new data source, allowing data engineers to exclude or down-weight polluted segments before model training begins—preventing the model from learning spam and adult content patterns in the first place.

  • Improvement: Use the hierarchical token dataset (630,684 records with categories and explanations) to build a real-time token-blocklist that is context-aware, not just keyword-based.

  • What the improved AI can do: When generating text, the AI can detect and suppress surface variants of polluted tokens (e.g., 老司机 or 威尼斯人) even when they appear in novel contexts, because the system understands the compositional nature of pollution (e.g., 菲律宾 + "申博" → gambling brand). This reduces harmful outputs without over-censoring normal uses of ambiguous words.

  • Improvement: Implement a periodic re-auditing mechanism that tracks the temporal evolution of pollution (as shown in the 2021–2026 Common Crawl analysis) and updates the model's internal safety filters accordingly.

  • What the improved AI can do: Automatically detect when new pollution families emerge (e.g., new gambling euphemisms) and when old ones fade, then adjust its generation behavior in real-time—preventing the model from being exploited by newly appearing spam patterns that weren't in its original training data.

  • Improvement: Use the token-sharing analysis (96.4% of polluted tokens appear in only one corpus) to build a pollution fingerprint for each corpus, enabling detection of contamination when combining multiple data sources.

  • What the improved AI can do: When a model is trained on multiple corpora, it can identify which corpus contributed which polluted tokens, allowing targeted re-filtering of specific data sources rather than re-training the entire model. It can also predict which new corpora are likely to introduce novel pollution based on similarity to known-polluted sources.

  • Improvement: Leverage the classification rationale and web-context fields in the released dataset to train a safety classifier that explains why a token is polluted, rather than just flagging it.

  • What the improved AI can do: When asked to generate or filter content, the AI can provide human-readable explanations for its safety decisions (e.g., this token is flagged because it's a known euphemism for adult content on Chinese gambling sites), improving transparency and allowing users to override false positives with confidence.

  • Improvement: Adopt the 148× speedup and 35.8× memory reduction methodology to build a lightweight, always-on monitoring system for live web data streams.

  • What the improved AI can do: Continuously audit new web data in near-real-time (reducing 1TB audits from months to hours), enabling rapid response to emerging pollution trends—such as the observed spike in adult content from 28% to 68% between 2024 and 2026—before they contaminate future model updates.

Abstract

Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three challenges: (1) their web-scale size makes full scan costly; (2) prior analyses are often too coarse to expose token-level pollution; (3) Chinese web pollution is implicit and rapidly changing. We propose Sampled-BPE, a lightweight token-level auditing pipeline that sample a small subset and train BPE tokenizer to surface polluted tokens. Experiments show that Sampled-BPE preserves usable estimates while substantially reducing runtime and memory: a 148.4 times speedup and a 35.8 times memory reduction induce only 4.25% relative error for pollution categories. We apply the pipeline to 11 open Chinese corpora and 6 Chinese Common Crawl snapshots from 2021 to 2026. The audit reveals widespread but uneven pollution across open corpora, as well as highly polluted and temporally shifting Chinese web content. We further release a hierarchical Chinese web token dataset with 660k+ token records, each with web context, category, and explanation fields, organized as trees to support review and tracing of pollution.

Sources

Related papers