How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
cs.CL, cs.LG
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/pangramlabs/WildAI
Terminology
Sources
- Pangram 4 Technical Report
- DeGenTWeb: A First Look at LLM-dominant Websites
- Relative Scaling Laws for LLMs
- Scaling Laws for Neural Language Models
- A Bitter Lesson for Data Filtering
- Bridging Compute- and Data-Optimal Pretraining
- StoryScope: Investigating idiosyncrasies in AI fiction
- Position: Model Collapse Does Not Mean What You Think
- Scaling Laws for Mixture Pretraining Under Data Constraints
- Rate of Model Collapse in Recursive Training
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering