Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT
cs.LG, cs.CL
Submitted: 2026-02-11
Updated: 2026-09-16
Code: https://github.com/project-numina/aimo-progress-prize
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models are commonly adapted to downstream tasks through supervised fine-tuning (SFT), but substantial distribution mismatch between downstream supervision and a model's generation
Terminology
Abstract
Large language models are commonly adapted to downstream tasks through supervised fine-tuning (SFT), but substantial distribution mismatch between downstream supervision and a model's generation distribution can intensify catastrophic forgetting. Data rewriting offers a data-centric way to narrow this mismatch before SFT. Existing methods, however, typically sample rewrites from a prompt-induced conditional distribution, which need not align with the backbone's natural question-answering generation distribution, and fixed templates can reduce output diversity. We formulate data rewriting as a policy-learning problem and train a lightweight LoRA rewriting policy with reinforcement learning. The policy optimizes question-answering-style distributional alignment and semantic diversity under a hard task-consistency gate, producing verified supervision for downstream SFT. Across three instruction-tuned backbones, the resulting models attain downstream gains broadly comparable to standard SFT while reducing degradation on non-downstream benchmarks in every evaluated setting. Additional experiments on logical reasoning and medical question answering provide preliminary evidence that a rewriting policy can be reused across domains for the same backbone.
Sources
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition
- The Llama 3 Herd of Models
- Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal
- Mistral 7B
- Self-Evolving LLMs via Continual Instruction Tuning
- Understanding Catastrophic Forgetting in Language Models via Implicit Inference
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
- GPT-4 Technical Report
- Preserving Diversity in Supervised Fine-Tuning of Large Language Models
- Let's Verify Step by Step
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- Gradient Episodic Memory for Continual Learning
- An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
- Training language models to follow instructions with human feedback
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks