Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards
cs.LG
Submitted: 2026-09-03
Updated: 2026-09-03
Code: https://github.com/LQgdwind/GAR
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories.
Terminology
Abstract
Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories. Existing dense reward alternatives, from surface heuristics to process reward models, either ignore the expert solutions already present in training corpora or require expensive offline annotation. We propose Gradient-Aligned Reward (GAR), which operates in the policy's own gradient space: truncated backpropagation through the output projection layer extracts a compact gradient vector for each rollout, and cosine similarity with an expert-anchor gradient yields a dense, reasoning-aware reward with less than 9% wall-clock overhead. We prove that this cosine admits a multiplicative decomposition into prediction-error and activation-pattern factors, providing a concrete characterization of what the alignment signal measures. On Qwen3-4B and Qwen3-8B, GAR consistently improves over GRPO and other baselines on competition-level math benchmarks and transfers to GPQA Diamond and MMLU-Pro without domain-specific data. Code and data are available at https://github.com/LQgdwind/GAR.
Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- Process Reward Models That Think
- Training Verifiers to Solve Math Word Problems
- Process Reinforcement through Implicit Rewards
- Towards Distillation-Resistant Large Language Models: An Information-Theoretic Perspective
- Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation
- Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
- MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
- Understanding R1-Zero-Like Training: A Critical Perspective
- Improve Mathematical Reasoning in Language Models by Automated Process Supervision
- Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
- TRAK: Attributing Model Behavior at Scale
- Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Learning Dynamics in RL Post-Training for Language Models
- Solving math word problems with process- and outcome-based feedback
- GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks