STRETCH the Boundaries: A Unified Self-Taught Framework for Progressive LLM Evolution
cs.CL
Submitted: 2026-09-16
Updated: 2026-09-16
Code: https://github.com/GuanNiPiShi123/STRETCH
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators
- G-Zero: Self-Play for Open-Ended Generation from Zero Data
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages
- GPT-4 Technical Report
- Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
- Advancing Reasoning in Large Language Models: Promising Methods and Approaches
- Guiding Reasoning in Small Language Models with LLM Assistance
- Evolving Deeper LLM Thinking
- Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
- AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
- The Llama 3 Herd of Models
- From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- COvolve: Adversarial Co-Evolution of Large-Language-Model-Generated Policies and Environments via Two-Player Zero-Sum Game
- LLMs and the ZPD
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering