StepOPSD: Step-Aware Online Preference Self-Distillation for Agent Reinforcement Learning
cs.AI
Submitted: 2026-05-26
Updated: 2026-08-26
Terminology
Sources
- A General Theoretical Paradigm to Understand Learning from Human Preferences
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Self-Distilled Agentic Reinforcement Learning
- Privileged Information Distillation for Language Models
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Policy Distillation
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Proximal Policy Optimization Algorithms
- Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
- Self-Distilled RLVR
- ReAct: Synergizing Reasoning and Acting in Language Models
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection