Extractable Memorization From First Principles
cs.LG, cs.CL
Submitted: 2026-07-14
Updated: 2026-09-25
Terminology
Sources
- Extracting books from production language models
- Extracting alignment data in open models
- The Files are in the Computer: On Copyright, Memorization, and Generative AI
- Report of the 1st Workshop on Generative AI and Law
- Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
- Extracting memorized pieces of (copyrighted) books from open-weight language models
- Estimating near-verbatim extraction risk in language models with decoding-constrained beam search
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- The Llama 3 Herd of Models
- Revisiting the LiRA Membership Inference Attack Under Realistic Assumptions
- Copyright Violations and Large Language Models
- Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply Chain
- Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
- LLM Dataset Inference: Did you train on my dataset?
- How much do language models memorize?
- Scalable Extraction of Training Data from (Production) Language Models
- Machine Learners Should Acknowledge the Legal Implications of Large Language Models as Personal Data
- 2 OLMo 2 Furious
- Rethinking LLM Memorization through the Lens of Adversarial Compression
- Gemma 2: Improving Open Language Models at a Practical Size
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks