Not All Preferences Deserve Gradients: Understanding Gradient Utility in Offline Reasoning Alignment
cs.AI
Submitted: 2026-02-01
Updated: 2026-09-03
Comments: EMNLP Main 2026
Code: https://github.com/hkust-nlp/simpleRL-reasonhttps:
License: http://creativecommons.org/licenses/by/4.0/
The gist: Offline preference optimization aligns reasoning models from fixed chosen--rejected pairs, yet standard methods apply gradient updates from every pair regardless of its training value under the
Terminology
Abstract
Offline preference optimization aligns reasoning models from fixed chosen--rejected pairs, yet standard methods apply gradient updates from every pair regardless of its training value under the current policy. We argue that this uniform treatment is wasteful and potentially harmful. From the perspective of gradient utility, we show that a pair's contribution depends jointly on informativeness and stability. Pair utility drifts as the policy evolves, high-gradient samples can coincide with high-curvature regions, leading to noisy and destabilizing updates, and the most effective supervision comes from stable confident errors where the model is reliably wrong yet curvature remains low. These findings motivate SAGE (Stability-Aware Gradient Efficiency), which maintains difficulty-stratified candidate pools refreshed during training and selects pairs within each pool by a forward-pass signal-to-curvature score. Only pairs with high current utility receive gradient computation; the rest are excluded from backpropagation. On mathematical reasoning benchmarks across multiple model scales, SAGE outperforms full-data and size-matched baselines while producing substantially smoother optimization trajectories.
Sources
- Training Language Models to Reason Efficiently
- Training Verifiers to Solve Math Word Problems
- KTO: Model Alignment as Prospect Theoretic Optimization
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
- DeepSeek-V3 Technical Report
- Qwen2.5 Technical Report
- OpenAI o1 System Card
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- MathScale: Scaling Instruction Tuning for Mathematical Reasoning
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection