Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
cs.LG, cs.AI
Submitted: 2026-05-19
Updated: 2026-09-09
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce a novel method that averages the logits of a frozen reference policy (e.g., SFT) and a trainable policy, and incorporate the method into Group Relative Policy Optimization (GRPO).
Terminology
Abstract
We introduce a novel method that averages the logits of a frozen reference policy (e.g., SFT) and a trainable policy, and incorporate the method into Group Relative Policy Optimization (GRPO). In contrast to Reinforcement Learning with Verifiable Rewards (RLVR) methods, our proposal does not involve a Kullback Leibler (KL) regularization or critic; the trainable policy and the reference anchor are coupled through the logit averaging structure to leverage the reasoning expertise of the trainable policy while maintaining the formatting advantage of SFT. Our method is evaluated on MATH, cn-k12, and MMLU, and the results show a higher accuracy or at least comparable accuracy relative to the canonical KL-regularized GRPO.
Sources
- Reinforcement Learning via Self-Distillation
- Tuning Language Models by Proxy
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Proximal Policy Optimization Algorithms
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks