Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model
cs.CL
Submitted: 2026-02-06
Updated: 2026-08-31
Comments: Accepted to ICML 2026. 53 pages, 14 figures, 22 tables. Code is publicly available at https://github.com/jacqueline-he/anchored-decoding
Code: https://github.com/jacqueline-he/anchored-decoding
License: http://creativecommons.org/licenses/by/4.0/
The gist: Language models (LMs) tend to memorize portions of their training data and emit verbatim spans.
Terminology
Abstract
Language models (LMs) tend to memorize portions of their training data and emit verbatim spans. When the underlying sources are sensitive or copyright-protected, such reproduction raises issues of consent and compensation for creators and compliance risks for developers. We propose Anchored Decoding, a plug-and-play inference-time method for suppressing verbatim copying: it enables decoding from any risky LM trained on mixed-license data by keeping generation in bounded proximity to a permissively trained safe LM. Anchored Decoding adaptively allocates a user-chosen information budget over the generation trajectory and enforces per-step constraints that yield a sequence-level guarantee, enabling a tunable risk-utility trade-off. To make Anchored Decoding practically useful, we introduce a new permissively trained safe model (TinyComma 1.8B), as well as Anchored Byte Decoding, a byte-level variant of our method that enables cross-vocabulary fusion via the ByteSampler framework (Hayase et al., 2025). Across six model pairs on long-form metrics for copying risk and utility, Anchored and Anchored Byte Decoding define a new Pareto frontier, preserving near-original fluency and factuality while closing up to 75% of the measurable copying gap between the risky baseline and a safe reference, at a modest inference overhead.
Sources
- The KL3M Data Project: Copyright-Clean Training Resources for Large Language Models
- Accelerating Large Language Model Decoding with Speculative Sampling
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Gemini: A Family of Highly Capable Multimodal Models
- Gemma 3 Technical Report
- CommonCanvas: An Open Diffusion Model Trained with Creative-Commons Images
- The Llama 3 Herd of Models
- Position: The Most Expensive Part of an LLM should be its Training Data
- Sampling from Your Language Model One Byte at a Time
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
- GPT-4 Technical Report
- On Provable Copyright Protection for Generative Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering