PACT: From Credit Assignment to Critic Alignment
cs.LG, cs.AI
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/AllSpark-Research/PACT
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted mathematical definition, leaving its relationship
Terminology
Abstract
Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted mathematical definition, leaving its relationship to commonly used training signals unclear. We formulate three regularity conditions, namely Completeness, Prefix Consistency, and Neutrality, and prove that they uniquely determine token-level credit. This characterization provides a unified basis for explaining phenomena across existing algorithms and guides the development of an improved actor-critic training procedure. Through this lens, an ideal teacher in On-Policy Distillation (OPD) acts as an implicit critic, yielding an expected policy gradient proportional to that induced by token-level credit. Response-level REINFORCE Leave-One-Out (RLOO) signals match the expected policy-gradient contribution of token-level credit despite their coarser granularity. We further establish approximate credit sparsity under bounded outcome rewards and show how intermediate critic errors in Generalized Advantage Estimation (GAE) can become comparable to the underlying credit. These motivate Policy Aligned Critic Training (PACT), which adopts an Actor-then-Critic update order to apply importance sampling correction to critic training and better align the critic with the updated policy. In agentic mathematical reasoning, PACT achieves 72.87% average accuracy across four benchmarks, outperforming GRPO and PPO by 8.80 and 13.16 percentage points, respectively. On SWE-bench Verified, PACT achieves a pass rate of 67.4%, outperforming PPO, GRPO, and SAO by 2.4, 2.0, and 3.8 percentage points, respectively.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
- Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
- daVinci-Env: Open SWE Environment Synthesis at Scale
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
- REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
- VinePPO: Refining Credit Assignment in RL Training of LLMs
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
- Understanding R1-Zero-Like Training: A Critical Perspective
- Counterfactual Credit Assignment in Model-Free Reinforcement Learning
- Would I have gotten that reward? Long-term credit assignment by counterfactual contribution analysis
- A Survey of Temporal Credit Assignment in Deep Reinforcement Learning
- Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks