Effective Synthetic Data Curation Requires Group-Level Signals
cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Terminology
Sources
- Phi-4 Technical Report
- Nemotron-4 340B Technical Report
- On the Diversity of Synthetic Data and its Impact on Training Large Language Models
- Strong Model Collapse
- The Vendi Score: A Diversity Evaluation Metric for Machine Learning
- The Llama 3 Herd of Models
- Studying Large Language Model Generalization with Influence Functions
- Datamodels: Predicting Predictions from Training Data
- DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models
- Kimi K2: Open Agentic Intelligence
- DataComp-LM: In search of the next generation of training sets for language models
- Textbooks Are All You Need II: phi-1.5 technical report
- DeepSeek-V3 Technical Report
- Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
- AgentInstruct: Toward Generative Teaching with Agentic Flows
- Recycling the Web: A Method to Enhance Pre-training Data Quality and Quantity for Language Models
- How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
- 2 OLMo 2 Furious
- Qwen2.5 Technical Report
- Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering