When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain
cs.CL
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
- Scaling Self-Play with Self-Guidance
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- Training Verifiers to Solve Math Word Problems
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Strong Model Collapse
- A Tale of Tails: Model Collapse as a Change of Scaling Laws
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
- Reinforced Self-Training (ReST) for Language Modeling
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- G-Zero: Self-Play for Open-Ended Generation from Zero Data
- Inference-Time Reward Hacking in Large Language Models
- R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training
- Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training
- Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL
- How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
- Can Large Reasoning Models Self-Train?
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering