QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning
cs.LG
Submitted: 2026-02-04
Updated: 2026-09-17
Code: https://github.com/volcengine/verl
License: http://creativecommons.org/licenses/by/4.0/
The gist: GRPO-style reinforcement learning (RL)-based LLM fine-tuning algorithms have recently gained popularity.
Terminology
Abstract
GRPO-style reinforcement learning (RL)-based LLM fine-tuning algorithms have recently gained popularity. Relying on heuristic trust-region approximations, however, they can lead to brittle optimization behavior, as global importance-ratio clipping and group-wise normalization fail to regulate samples whose importance ratios fall outside the clipping range. We propose Query-Adaptive Trust-Region policy Optimization (QUATRO), which directly enforces trust-region constraints through a principled optimization. This yields a clear and interpretable objective that enables explicit control over policy updates and stable, entropy-controlled optimization, with a stabilizer terms arising intrinsically from the exact trust-region formulation. Empirically verified on diverse mathematical reasoning benchmarks, QUATRO shows stable training under increased policy staleness and aggressive learning rates, maintaining well-controlled entropy throughout training.
Sources
- LSPO: Length-aware Dynamic Sampling for Policy Optimization in LLM Reasoning
- Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
- Reasoning with Exploration: An Entropy Perspective
- GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning
- Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
- XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation
- AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- Trust Region Constrained Measure Transport in Path Space for Stochastic Optimal Control and Inference
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- Evaluating Large Language Models Trained on Code
- Measuring Mathematical Problem Solving With the MATH Dataset
- SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
- RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
- Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
- Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks