Self-Play Search Distillation for Large Language Model Reasoning
cs.AI
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- Policy Gradient Search: Online Planning and Expert Iteration without Search Trees
- Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
- Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning
- GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
- Stream of Search (SoS): Learning to Search in Language
- TextArena
- Can Large Language Models Play Games? A Case Study of A Self-Play Approach
- RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
- Distilling the Knowledge in a Neural Network
- lmgame-Bench: How Good are LLMs at Playing Games?
- ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context
- SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
- OpenSIR: Open-Ended Self-Improving Reasoner
- OpenSpiel: A Framework for Reinforcement Learning in Games
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
- SPADE: Self-Play in Adaptive Synthetic Executable Environments
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- Amortized Planning with Large-Scale Transformers: A Case Study on Chess
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Self-Distillation Enables Continual Learning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection