Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT

arXiv:2602.11220 · cs.LG, cs.CL · Submitted 2026-02-11 · Read on arXiv

cs.LG, cs.CL

Submitted: 2026-02-11

Updated: 2026-09-16

Code: https://github.com/project-numina/aimo-progress-prize

License: http://creativecommons.org/licenses/by/4.0/

The gist: Large language models are commonly adapted to downstream tasks through supervised fine-tuning (SFT), but substantial distribution mismatch between downstream supervision and a model's generation

Terminology

Abstract

Large language models are commonly adapted to downstream tasks through supervised fine-tuning (SFT), but substantial distribution mismatch between downstream supervision and a model's generation distribution can intensify catastrophic forgetting. Data rewriting offers a data-centric way to narrow this mismatch before SFT. Existing methods, however, typically sample rewrites from a prompt-induced conditional distribution, which need not align with the backbone's natural question-answering generation distribution, and fixed templates can reduce output diversity. We formulate data rewriting as a policy-learning problem and train a lightweight LoRA rewriting policy with reinforcement learning. The policy optimizes question-answering-style distributional alignment and semantic diversity under a hard task-consistency gate, producing verified supervision for downstream SFT. Across three instruction-tuned backbones, the resulting models attain downstream gains broadly comparable to standard SFT while reducing degradation on non-downstream benchmarks in every evaluated setting. Additional experiments on logical reasoning and medical question answering provide preliminary evidence that a rewriting policy can be reused across domains for the same backbone.

Sources

Related papers