Multi-Turn On-Policy Distillation with Prefix Replay
Baohao Liao, Hanze Dong, Christof Monz, Xinxing Xu, Li Dong, Furu Wei
cs.LG, cs.AI, cs.CL, stat.ML
Submitted: 2026-07-06
Code: https://github.com/THUDM/slime
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
- ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent
- A General Theoretical Paradigm to Understand Learning from Human Preferences
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- FireAct: Toward Language Agent Fine-tuning
- Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
- RLHF Workflow: From Reward Modeling to Online RLHF
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
- Reinforced Self-Training (ReST) for Language Modeling
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Measuring Mathematical Problem Solving With the MATH Dataset
- Distilling the Knowledge in a Neural Network
- OpenAI o1 System Card
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
- Self-Hinting Language Models Enhance Reinforcement Learning
- Statistical Rejection Sampling Improves Preference Optimization
- Orca-Math: Unlocking the potential of SLMs in Grade School Math
- WebGPT: Browser-assisted question-answering with human feedback
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks