AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning".
Jane: The paper was written by Can Jin, Yang Zhou, Qixin Zhang, Hongwu Peng, Di Zhang et al. from Rutgers University and Nanyang Technological University and Fudan University and University of Connecticut and Red Hat AI Innovation and MIT-IBM Watson AI Lab and Massachusetts Institute of Technology (MIT) and NVIDIA Research.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary and Core Insight: Tom: The paper summarizes that current methods for improving LLM reasoning—Test-Time Scaling or TTS—are either unstable, like pure RL, or they are static and brittle because of PRMs.
Jane: They found a sweet spot in between those two approaches by leveraging AIRL and GRPO to create something called the AIRL-S framework.
Lu: I love the idea that this isn's not just a patch; it's an architectural change, fundamentally redefining how we build and train reasoning capabilities into the AI.
Meng: The core insight they developed is that your reward function learned during training is actually the best Process Reward Model for search at inference time, which seems like a massive efficiency win.
Lalam: It suggests that all those expensive process labels we usually need to teach an AI how to think step-by-step, we might not need them at all.
Improvements and Results: Tom: Let's talk about the results, because they're pretty impressive—a nine percent average improvement over the base model across eight different reasoning tasks.
Jane: It’s not just a boost; it’s that our PRM consistently outperforms all the other static PRMs trained on labeled data, which is huge for robustness.
Lu: This isn't just academic success; this is proof that we can fundamentally change the trajectory of how hard problems are solved by machines.
Meng: Matching GPT-4o’s performance while using a more cost-effective, generalized PRM is exactly what I want to see in deployment, reducing the operational overhead dramatically.
Lalam: It’s about creating reliable reasoning engines that can generalize well enough to handle complex tasks without needing constant retraining or fine-tuning.
Conclusion and Wrap Up: Tom: So, we've seen how AIRL-S unifies RL and Search methods, moving beyond the old limitations of both approaches.
Jane: It’s a robust framework that allows us to scale inference computation without relying on massive amounts of labeled data.
Lu: We are seeing the future where the AI doesn's just output an answer, but it can guide its own path to the a correct one, step by step.
Meng: The engineering implication is that we can start building these systems more reliably and integrate them into real-world applications sooner.
Lalam: We are achieving a new level of reasoning maturity with this unified approach.
Final Thoughts on Impact: Tom: As we wrap up, I want to hear from each of you about the bigger picture—what does this all mean for the world?
Jane: From a practical standpoint, it means better tutors and better verification tools are becoming much more accessible.
Lu: I think this is how we accelerate scientific discovery; we're giving AI the ability to hypothesize and verify complex chains of logic at an unprecedentedly high speed.
Meng: My concern is that this opens up a new kind of opportunity for software reliability, where the PRM can be used to verify code generation quality before deployment.
Lalam: I see a future where this technology enables AI to assist in decision-making processes in high-stakes environments, offering truly reliable guidance based on its learned understanding of the process.
Can Jin, Yang Zhou, Qixin Zhang, Hongwu Peng, Di Zhang, Zihan Dong, Marco Pavone, Ligong Han (Nanyang Technological University), Zhang-Wei Hong (Massachusetts Institute of Technology), Tong Che (NVIDIA Research), Dimitris N. Metaxas, Can Jin, Yang Zhou, Qixin Zhang, Hongwu Peng, Di Zhang, Zihan Dong, Marco Pavone, Ligong Han (Nanyang Technological University), Zhang-Wei Hong (Massachusetts Institute of Technology), Tong Che (NVIDIA Research), Dimitris N. Metaxas
Rutgers University · Nanyang Technological University · Fudan University · University of Connecticut · Red Hat AI Innovation · MIT-IBM Watson AI Lab · Massachusetts Institute of Technology (MIT) · NVIDIA Research
cs.LG, cs.AI
Submitted: 2025-08-19
Updated: 2026-08-25
Code: https://github.com/deepseek-ai/DeepSeek-Math
Project page: https://livecodebench.github.io
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
The gist: " * Problem Statement and Motivation Test-time scaling (TTS) aims to enhance the reasoning performance of Large Language Models (LLMs) by allocating additional inference-time computation.
Key concepts
- AIRL-S
- This framework unifies Reinforcement Learning and search methods for LLM reasoning. It addresses limitations of existing Test-Time Scaling (TTS) approaches by leveraging Adversarial Inverse Reinforcement Learning (AIRL) to create a robust system.
- Process Reward Model (PRM)
- A PRM is a model that evaluates the quality of a process or sequence of steps taken by an AI. The paper suggests that the reward function learned during training can serve as an effective, generalized PRM for inference time.
- Test-Time Scaling (TTS)
- TTS refers to current methods used to improve LLM reasoning capabilities. The episode notes that these existing methods are either unstable (like pure RL) or static and brittle (due to PRMs).
- Adversarial Inverse Reinforcement Learning (AIRL)
- AIRL is a technique leveraged by the AIRL-S framework. It helps create the necessary reward function, allowing the AI to develop reliable reasoning engines without needing massive amounts of labeled data.
Terminology
Summary
"
Problem Statement and Motivation
Test-time scaling (TTS) aims to enhance the reasoning performance of Large Language Models (LLMs) by allocating additional inference-time computation. Traditionally, two distinct paradigms have existed:
-
Reinforcement Learning (RL) methods: These optimize sparse outcome-based rewards but
suffer from instability and low sample efficiency.
-
Search-based techniques: These are guided by independently trained, static process reward models (PRMs). However, these require
expensive human- or LLM-generated labels
and oftendegrade under distribution shifts.
The core motivation of this research is to address these limitations by investigating how to effectively combine RL-based and Search-based TTS. The paper posits a central insight: the reward function learned during RL training inherently represents the ideal PRM for guiding downstream search.
Proposed Solution: AIRL-S
The authors introduce AIRL-S, a framework that provides natural unification of RL-based and search-based TTS.
The system is designed to eliminate the need for labeled intermediate process data by leveraging two key components: Adversarial Inverse Reinforcement Learning (AIRL) and Group Relative Policy Optimization (GRPO).
Methodology
- ** Data Generation:**
The training dataset consists of questions q in Q. From the reference rollouts, a Chain-of-Thought (CoT), C(q) = C 1, C 2,, C T, is obtained. A binary outcome reward indicates whether a CoT yields the correct final answer.
- ** Learning the Process Reward Model (PRM) via AIRL:**
To avoid costly step-wise labels, AIRL is used to train a discriminator D phi that distinguishes the reference rollouts from the policy outputs.
-
State and Action: The state is defined as
the question q together with the preceding reasoning steps C 1,, C i-1, and action: the current step C i.
-
Discriminator Definition: The discriminator is defined as:
D phi (C i q, C<i) = f phi (q, C i) over f phi (q, C i) + pi theta(C i q, C<i)
- Step-wise Reward: The step-wise reward r phi is derived from this discriminator:
r phi (C i q, C<i) = D phi (C i q, C<i) over 1 - D phi (C i q, C<i)
The PRM is trained by minimizing the LAIRL loss:
LAIRL = sum X about Q, C about pi e D phi(C q, C<i) - E X about Q, C about pi theta 1 - D phi(C q, C<i)
This loss is used to update the discriminator D phi, which serves as the reward function r phi.
- ** Policy Training via Combined Objectives:**
The policy model pi theta is optimized using a composite objective that combines AIRL and GRPO:
J(theta) = lambda J AIRL(theta) + (1 - lambda) J GRPO(theta)
-
AIRL Objective (J AIRL): This aims to maximize the discriminator-derived rewards, essentially
fool[ing] the discriminator into classifying its rollouts are reference rollouts.
-
GRPO Objective (J GRPO): This uses binary outcome rewards (based on whether a group of CoTs yields the correct final answer) and is defined as:
J GRPO(theta) = E X about Q, C k k=1 G about pi old(timesq) [#C k q over# G - DKL(pi theta pi ref) - A k]
The policy is updated by maximizing this composite objective, J(theta).
- ** PRM-guided Test-Time Search:**
At inference time, the learned r phi guides various search algorithms:
-
Best-of-N Sampling: Samples N full solutions and selects aggregate the best using a step-wise aggregation score.
-
Beam Search: Extends each beam node by M candidate steps, ranking them by r phi, and retaining the top N nodes.
-
Monte Carlo Tree Search (MCTS): Uses UCT (Upper Confidence Bound for Trees) criteria, combining PRM rewards with visit counts to select nodes.
A key aggregation method is PRMMin-Sum:
-
The step-wise aggregation score for a chain C k is defined as the minimum PRM reward among its steps: s C k = i r phi C ik..
-
Inter-answer aggregation is performed by aggregating the step-wise scores of all solutions that yield the same distinct answer, S(a) = sum scores for answer a. The final prediction is the answer with the largest S(a).
Experimental Results and Findings
The experiments were conducted on eight standard reasoning benchmarks (mathematics, science, and code generation).
-
Policy Performance: The Qwen2.5-7B-AIRL-S policy model achieved an average improvement of 9% over the base model, matching the performance of GPT-4o.
-
PRM Generalizability: The Qwen2.5-AIRL-S-PRM consistently outperformed PRMs trained on labeled data (e.g., outperforming Math-Shepherd by 2.4%).
-
Ablation Study: Adding the PRM significantly improved both training and validation performance, enabling longer response generation at test time compared to GRPO trained with outcome rewards only.
Conclusion
The research concludes that AIRL-S provides a robust and cost-effective solution to complex reasoning tasks in LLMs,
validating that your reward function for RL is your best PRM for search.
Improvements for AI systems
The following improvements detail the implementation of the AIRL-S framework, transforming existing AI systems by integrating learned process rewards into both training and inference phases.
A. Implementation of Adversarial Inverse Reinforcement Learning (AIRL) for PRM Generation:
Instead of relying on sparse, binary outcome rewards (which lead to instability), the system will utilize an adversarial training loop to generate a dense, step-wise Process Reward Model (r phi). This requires training a discriminator D phi that distinguishes between reference rollouts (pi e, drawn from the replay buffer) and policy outputs (pi theta).
-
Specific Action: Implement the LAIRL loss function (Equation 2 in the paper). The system learns to maximize this adversarial objective, effectively training D phi to represent a high-quality, step-wise reward function r phi that guides its own policy.
-
System Capability:
-
Automated PRM Learning: The system generates a generalizable PRM without requiring any expensive human or LLM-generated labels for intermediate reasoning steps.
-
Efficiency: It drastically reduces the cost of building fine-grained reward models, mitigating the need for massive, labeled step-wise datasets.
B. Optimization via Combined Objective Functions (AIRL + GRPO):
The policy model (pi theta) will be optimized using a composite objective function that integrates the dense rewards from AIRL and the binary outcome rewards from Group Relative Policy Optimization (GRPO).
-
Specific Action: Implement the weighted composite objective J(theta) = lambda J AIRL(theta) + (1 - lambda) J GRPO(theta) (Equation 6), where lambda is a hyperparameter balancing process guidance and final accuracy.
-
System Capability:
-
Stable Refinement: The system achieves fine-grained, step-by-step improvement (via J AIRL) while maintaining the overall goal of task completion (via J GRPO), addressing the instability inherent in sparse reward systems.
A. Utilizing the Learned PRM (r phi) as a Heuristic Verifier:
The trained, policy-independent PRM (r phi) is deployed at inference time to guide and score candidate reasoning paths across various search algorithms.
-
Specific Action: Use r phi as the scoring function for all subsequent search procedures (Best-of-N, Beam Search, MCTS).
-
System Capability:
-
Guided Exploration: The system uses the PRM to score and retain only those intermediate steps that are highly probable/correct according to the learned reward structure, drastically pruning irrelevant reasoning paths.
-
Robustness: Since r phi is policy-independent (unlike static PRMs), it remains robust against distribution shifts when used across different LLMs or datasets.
B. Implementing Advanced Search Strategies Guided by r phi: The system will leverage the dynamic PRM within specific, advanced search paradigms:
-
Best-of-N / Beam Search: Candidates are scored using r phi. The system retains the top candidates based on these scores before expanding or finalizing paths.
-
Monte Carlo Tree Search (MCTS): The PRM is used to estimate the cumulative reward (mu n) for future steps, allowing MCTS to select nodes via the Upper Confidence Bound (UCT) criterion, prioritizing steps that are highly rewarded by r phi.
-
System Capability:
-
Enhanced Reasoning: The system moves beyond simple sampling or majority voting; it selectively explores and prioritizes paths based on step-wise quality, leading to superior reasoning chains.
A. Aggregation Strategy Implementation:
The system will implement various PRM aggregation methods (e.g., PRM-Min-Sum, PRM-Last-Max) to select the final answer from the set of N candidate solutions generated during search.
-
Specific Action: Use the calculated step-wise scores derived from r phi to aggregate results across chains that yield the same final answer.
-
System Capability:
-
Optimal Selection: The system selects the most reliable answer by synthesizing evidence from multiple rollouts, rather than merely picking a single path or relying on majority vote.
B. Cross-Task Generalization:
The architecture is designed to produce a unified PRM that serves as a drop-in verifier
across diverse tasks.
-
Specific Action: Train the AIRL-S framework on diverse datasets (e.g., mathematical, coding, scientific reasoning) allowing the r phi to generalize across these domains.
-
System Capability:
-
Versatility: The system can apply a single, learned PRM to guide search and evaluate performance across entirely different domains (e.g., evaluating a math problem using the same learned reward structure used for code generation).
Abstract
Test-time scaling strategies for Large Language Models predominantly rely on either reinforcement learning with sparse outcome rewards or search-based methods guided by static Process Reward Models. However, outcome-based RL often suffers from training instability and sample inefficiency, while static PRMs require expensive step-wise supervision and are susceptible to reward hacking due to distributional shifts. In this paper, we introduce AIRL-S, a unified framework that integrates Adversarial Inverse Reinforcement Learning with Group Relative Policy Optimization. By inferring a dense, step-wise reward model directly from reference trajectories, AIRL-S eliminates the dependency on labeled process data and uses the same learned PRM as both a training signal and a verifier for search-based TTS. Extensive evaluations across eight benchmarks in mathematics, science, and code generation demonstrate that our policy model improves average performance by 9% over the base model, matching GPT-4o. We further analyze how the AIRL and GRPO objectives complement each other and how the learned PRM transfers across generators and search algorithms, establishing a robust and cost-effective methodology for scaling test-time computation in complex reasoning tasks.
Sources
- Phi-4 Technical Report
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- A Performance Study of LLM-Generated Code on Leetcode
- Process Reinforcement through Implicit Rewards
- A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Coding Challenge Competence With APPS
- Qwen2.5-Coder Technical Report
- Rewarding Chatbots for Real-World Engagement with Millions of Users
- Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning
- Disentangling Memory and Reasoning Ability in Large Language Models
- TACO: Topics in Algorithmic COde generation dataset
- DeepSeek-V3 Technical Report
- GuardReasoner: Towards Reasoning-based LLM Safeguards
- GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
- s1: Simple test-time scaling
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks
- Asymptotic Optimality of Thompson Sampling for Risk-Averse Bandits with Sub-Gaussian Rewards