A Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling Exponents
cs.LG, cs.AI
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/alexandrumeterez/QuadraticModel
Terminology
Sources
- A Latent Variable Model Approach to PMI-based Word Embeddings
- Explaining Neural Scaling Laws
- On the origin of neural scaling laws: from random graphs to natural language
- Scaling Trends for Data Poisoning in LLMs
- Towards a theory of how the structure of language is acquired by deep neural networks
- How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Model
- Deriving Neural Scaling Laws from the statistics of natural language
- Poisoning Web-Scale Training Datasets is Practical
- Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Learning Curve Theory
- Scaling Laws for Neural Language Models
- Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models
- Symmetry in language statistics shapes the geometry of model representations
- Stronger Data Poisoning Attacks Break Data Sanitization Defenses
- On the Emergence of Linear Analogies in Word Embeddings
- Universal One-third Time Scaling in Learning Peaked Distributions
- Safety Pretraining: Toward the Next Generation of Safe AI
- A Solvable Model of Neural Scaling Laws
- Deep networks learn to parse uniform-depth context-free languages from local statistics
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks