Boosting LLM Exploration via Weak-Model Guidance in RLVR
cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
Terminology
Sources
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation
- Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
- Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
- RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Beyond the Sampled Token: Preserving Candidate Support in RLVR
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Decomposed Prompting: A Modular Approach for Solving Complex Tasks
- OpenAI GPT-5 System Card
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
- Automatic Chain of Thought Prompting in Large Language Models
- The Invisible Leash: Why RLVR May or May Not Escape Its Origin
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering