It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs
cs.CL, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/xjdr-alt/entropix
Project page: https://physics.allen-zhu.com
Terminology
Sources
- Phi-4 Technical Report
- Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
- LFM2 Technical Report
- The KL3M Data Project: Copyright-Clean Training Resources for Large Language Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Extracting memorized pieces of (copyrighted) books from open-weight language models
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- OpenThoughts: Data Recipes for Reasoning Models
- Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
- propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- Data-efficient pre-training by scaling synthetic megadocs
- Kimi K2.5: Visual Agentic Intelligence
- Kimi K2: Open Agentic Intelligence
- WebSailor: Navigating Super-human Reasoning for Web Agent
- Combating Data Laundering in LLM Training
- Textbooks Are All You Need II: phi-1.5 technical report
- Learning Facts at Scale with Active Reading
- Best Practices and Lessons Learned on Synthetic Data
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering