Self-Improving Large Language Models via Progressive Experience Evolution
Shijie Ren, Xiting Wang, Meng Li, Yujie Guo, Yunhang Yao, Ziheng Peng, Xunlong Wang, Yuetan Chen, Haoyang Zhou, Yunlong Liang, Fandong Meng
cs.CL, cs.AI, cs.LG
Submitted: 2026-08-03
Comments: 10 pages, 5 figures
Code: https://github.com/rrrsj/SPEE
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Training Verifiers to Solve Math Word Problems
- MiniLLM: On-Policy Distillation of Large Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Distilling the Knowledge in a Neural Network
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Scaling Laws for Neural Language Models
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
- Let's Verify Step by Step
- ReFT: Reasoning with Reinforced Fine-Tuning
- Privileged Information Distillation for Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning
- Learning to summarize from human feedback
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering