Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?
Tsinghua University · Alibaba Group · Alibaba Group
cs.CL
Submitted: 2026-08-11
Updated: 2026-08-31
Code: https://github.com/qingjiesjtu/QGDE
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper "Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?" investigates whether released tokenizer vocabularies can be used to estimate the corpus ratios of
Terminology
Summary
The paper Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?
investigates whether released tokenizer vocabularies can be used to estimate the corpus ratios of individual tokens in hidden pretraining corpora. The authors first show that BPE tokenizers trained on different corpora share stable token ID–ratio distributions,
which motivates transferring this distribution from known corpora to a target tokenizer trained on hidden corpora. They then propose Quantile-Guided Density Estimation (QGDE),
which approximates this distribution with multiple quantile trends and uses local density weighting to produce token-level estimates.
In controlled settings and a realistic setting using the released SmolLM tokenizer, QGDE achieves mean relative errors as low as 3.00% for token-level estimation and 3.08% after aggregation into category-level mixtures.
The authors conclude that released tokenizer vocabularies provide a useful signal for fine-grained corpus estimation beyond coarse composition inference.
The code is available at https://github.com/qingjiesjtu/QGDE.
Improvements for AI systems
Improvements to AI Systems:
- Enhanced Data Auditing for Pretraining Corpora
-
Improvement: Integrate QGDE into AI training pipelines to estimate token-level composition of undisclosed or proprietary datasets used by third-party models.
-
Capability: The system can now audit whether a model’s pretraining data over-represents certain domains (e.g., medical text, code) or under-represents others, enabling compliance checks, bias detection, and informed fine-tuning decisions without access to the raw corpus.
- Adaptive Tokenizer Selection for Domain-Specific Fine-Tuning
-
Improvement: Use QGDE’s token-ratio estimates to compare a target tokenizer’s vocabulary against known domain distributions, then automatically recommend the most suitable tokenizer or vocabulary extension for a new task.
-
Capability: The AI can proactively suggest adding domain-specific tokens (e.g., legal jargon, scientific symbols) by estimating which tokens are missing or over/under-represented in the hidden corpus, improving downstream task performance and reducing tokenization fragmentation.
- Privacy-Preserving Data Mixture Optimization
-
Improvement: Apply QGDE to estimate the composition of a hidden corpus (e.g., a competitor’s model) and then adjust the training mixture of a new model to match or intentionally diverge from that distribution.
-
Capability: The system can replicate the strengths of a closed-source model’s training data (e.g., 70% code, 30% natural language) or deliberately avoid overfitting to a known distribution, all without ever accessing the original data—useful for synthetic data generation or targeted curriculum learning.
- Real-Time Token-Level Drift Monitoring in Continual Learning
-
Improvement: Deploy QGDE as a lightweight monitoring module that estimates token-ratio shifts in streaming data during online learning.
-
Capability: The AI can detect when the incoming data distribution deviates from the original training distribution (e.g., sudden increase in rare tokens), triggering automatic re-weighting, vocabulary updates, or rollback to prevent catastrophic forgetting or performance degradation.
- Cross-Model Vocabulary Alignment for Ensemble and Transfer Learning
-
Improvement: Use QGDE’s stable token-ID–ratio mapping to align vocabularies of different models trained on hidden corpora, enabling direct token-level probability fusion or knowledge distillation.
-
Capability: The system can combine outputs from multiple black-box models (e.g., different LLMs) by estimating their token-level confidence on a shared vocabulary, improving ensemble accuracy and enabling more reliable transfer of knowledge from a teacher to a student model without needing aligned training data.
- Cost-Effective Data Valuation and Licensing
-
Improvement: Leverage QGDE to estimate the proportion of specific token categories (e.g., copyrighted text, personal data) in a hidden corpus, aiding in legal risk assessment.
-
Capability: The AI can flag potential copyright or privacy violations by estimating the ratio of problematic tokens (e.g., verbatim quotes, PII patterns) in a model’s training data, allowing developers to make informed decisions about data sourcing or model release.
Abstract
Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model weights are released. Prior work has inferred corpus mixtures or traced specific token groups from released tokenizer vocabularies; in contrast, we estimate corpus ratios for arbitrary target tokens. We first show that BPE tokenizers trained on different corpora share stable token ID--ratio distributions, motivating distribution transfer from known corpora to a target tokenizer trained on hidden corpora. We then propose Quantile-Guided Density Estimation (QGDE), which approximates this distribution with multiple quantile trends and uses local density weighting to produce token-level estimates. In controlled settings and a realistic setting using the released SmolLM tokenizer, QGDE achieves mean relative errors as low as 3.00% for token-level estimation and 3.08% after aggregation into category-level mixtures. These results suggest that released tokenizer vocabularies provide a useful signal for fine-grained corpus estimation beyond coarse composition inference.
Sources
- Qwen Technical Report
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- The Foundation Model Transparency Index
- Training Compute-Optimal Large Language Models
- DeepSeek-V3 Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- How Does Code Pretraining Affect Language Model Task Performance?
- OpenAI GPT-5 System Card
- Benchmark Data Contamination of Large Language Models: A Survey
- Qwen3 Technical Report
- Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering