Self-Play Pretraining with Zero Data
cs.AI, cs.CL
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/amorehead/jvp_flash_attention
Terminology
Sources
- Scaling Self-Play with Self-Guidance
- Learning to See by Looking at Noise
- Self-Questioning Language Models
- Olmix: A Framework for Data Mixing Throughout LM Development
- Anchored Self-Play for Code Repair
- Neural Networks and the Chomsky Hierarchy
- STP: Self-play LLM Theorem Provers with Iterative Conjecturing and Proving
- Formal Theorem Proving by Rewarding LLMs to Decompose Proofs Hierarchically
- From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
- Learning Universal Predictors
- Neural Turing Machines
- Scaling Laws for Autoregressive Generative Modeling
- Training Compute-Optimal Large Language Models
- A Theory of Universal Artificial Intelligence based on Algorithmic Complexity
- Simple and Scalable Strategies to Continually Pre-train Large Language Models
- Scaling Laws for Neural Language Models
- Pre-training without Natural Images
- Data-efficient pre-training by scaling synthetic megadocs
- Training Language Models via Neural Cellular Automata
- DataComp-LM: In search of the next generation of training sets for language models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection