Semifactual Credit-Augmented Policy Optimization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Semifactual Credit-Augmented Policy Optimization".
Tom: Reinforcement learning with verifiable rewards (RLVR) has advanced large language model reasoning, but its predictions remain sensitive to task-irrelevant prompt features.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, to recap where we are, we’ve seen that this paper introduces a method called "Semifactual Credit-Augmented Policy Optimization," and I want to talk about what that title actually means in plain language for our listeners. Basically, they're taking the idea of rewarding good outcomes in reinforcement learning and adding a special layer that looks at how stable the model's predictions are when you slightly mess with the input prompt.
Jane: Exactly; think of it like this: instead of just giving a big reward or penalty based on whether the final answer is right or wrong, they are also giving credit to individual words in the response based on how much those words change when the prompt is tweaked in irrelevant ways. It’s about rewarding stability rather than just correctness, which is a clever distinction.
Lu: The authors are Zhejiang University and Westlake University and Shanghai Innovation Institute and Shanghai AI Laboratory, so it shows a strong collaboration across different major research hubs focused on this kind of deep sensitivity analysis in LLMs.
Meng: That kind of multi-institutional effort suggests the underlying problem is significant enough to warrant attention from several top research groups simultaneously, which points toward this being more than just an incremental tweak to an existing RL technique.
Lalam: For our culture here, this paper suggests we can move towards training systems that are less brittle; if we can build models that ignore irrelevant prompt noise, the resulting AI will be much more trustworthy and less prone to unpredictable behavior in production environments.
The paper's summary: Tom: Now let's look at what they actually found in the "Semifactual Credit-Augmented Policy Optimization" paper. They investigated how sensitive token predictions are by changing the prompt while keeping the problem and answer constant, and they showed there is a real variation in this sensitivity across different tokens within a response.
Jane: What that means is that some parts of the generated text are much more sensitive to prompt changes than others, which helps them build a signal for credit assignment. They then create a relative stability score, 'u', based on these drifts, and they use the negative part of this score to adjust the advantage given to tokens during training.
Lu: That token-level analysis is key because it moves away from just looking at the final answer; it’s about understanding the causal invariance of specific tokens under semifactual prompt interventions, which is a deeper investigation into how LLMs process information.
Meng: So, they are essentially using these measured drifts to filter out tokens that are likely just artifacts of prompt sensitivity during the early training phase, which is a very practical way to guide the model's learning trajectory without needing massive retraining cycles.
Lalam: This sounds like a way to build better internal representation; if we can train the AI to favor responses built on more stable token pathways, it creates a more coherent and less fragile reasoning structure within the model itself.
The paper's improvements: Tom: The main improvement they propose is this new way of constructing the token-level advantage during policy optimization. They modify the standard Group Relative Policy Optimization, or GRPO, by adding that semifactual credit signal to create a final token-level advantage equation.
Jane: The math shows that you take the original outcome-derived advantage from GRPO and then add a term based on the normalized stability score 'sgu−i,t', meaning unstable tokens get a lower advantage because of this augmentation, but only negatively.
Lu: By doing this, they ensure that relatively unstable tokens receive reduced advantages during the initial training phase, while explicitly stating that stability alone doesn't grant extra credit; it's a refinement mechanism for existing outcomes.
Meng: The authors use parameters like lambda zero and N zero to control how strong and how long this augmentation lasts during the start of training before they switch back to the standard GRPO process, which gives them fine-grained control over this process.
Lalam: This is powerful because it means we are actively teaching the AI *which* parts of its reasoning chain are reliable based on sensitivity analysis, rather than relying solely on the final reward signal.
Conclusion: Tom: So, wrapping up the discussion on "Semifactual Credit-Augmented Policy Optimization," this paper shows that incorporating semifactual stability provides a way to fine-tune credit assignment at the token level in RLVR. The results on models like Qwen3-4B-Base were quite encouraging, showing accuracy improvements over GRPO by five point six three and four point one seven percentage points on specific mathematics benchmarks, achieving top results across many evaluated methods for those scales.
Jane: It’s clear that this signal acts as an effective training signal for improving reasoning and generalization through this finer-grained credit assignment in RLVR, which is a really important finding for how we view model improvement.
Lu: The implications are that we can suppress high-drift token candidates during decoding to improve accuracy without updating the model weights, which suggests a way to refine the internal process of inference itself based on these stability metrics.
Meng: And I see the practical value in the ablation studies too; they confirmed that using negative-only corrections yields the highest average accuracy across competition-level benchmarks, supporting that principle of selectively reducing credit for relatively unstable tokens without rewarding stability alone.
Lalam: Ultimately, this work demonstrates how semifactual stability complements outcome rewards with an effective signal for finer-grained credit assignment in RLVR, and it gives us a much more nuanced tool to refine how we teach the AI to reason.
Junshu Pan, Zhizhang Fu, Shulin Huang, Yiran Ding, Zifan Cheng, Wenqi Shao
Zhejiang University · Westlake University
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/DtYXs/SCAPO
Importance score: 92/100
The gist: Reinforcement learning with verifiable rewards (RLVR) has advanced large language model reasoning, but its predictions remain sensitive to task-irrelevant prompt features.
Key concepts
- Semifactual Prompt Interventions
- This involves slightly altering task-irrelevant parts of a prompt, such as paraphrasing or adding minor typos, while keeping the core mathematical problem and its correct answer unchanged. This tests how sensitive a model's token predictions are to these irrelevant changes.
- Probability Drift Measurement
- A diagnostic setup measures the difference in probability for each response token between the original prompt and all perturbed prompts. This difference is quantified using a bounded-symmetric distance, which normalizes the change to provide a consistent metric for sensitivity across different tokens.
- Relative Stability Signal (u)
- This signal aggregates token-level drifts into a group score that measures how stable responses are relative to each other within a prompt group. Crucially, only the negative part of this score is kept, meaning the method penalizes unstable tokens without granting extra credit just for being stable.
- Semifactual Credit Augmentation
- SCAPO modifies standard policy optimization by adding a credit term based on the stability signal. This ensures that relatively unstable tokens receive lower advantages during early training phases, effectively guiding the model toward more robust reasoning pathways.
Terminology
Summary
Reinforcement learning with verifiable rewards (RLVR) has advanced large language model reasoning, but its predictions remain sensitive to task-irrelevant prompt features. This study introduces Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of Group Relative Policy Optimization (GRPO) that incorporates semifactual stability into token-level credit assignment to improve reasoning accuracy without updating model weights.
The gist
SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone.
How it works: Semifactual Prompt Interventions and Drift Measurement
The research investigates spurious dependence at the token level by using semifactual prompt interventions that preserve the underlying problem and answer. This involves altering task-irrelevant prompt features while keeping the mathematical problem and its answer fixed, such as paraphrasing, adding minor typo noise, or appending irrelevant context. To quantify sensitivity before RL training, a diagnostic setup is used where frozen models generate responses under original prompts and all perturbed prompts. The probability drift between the original and perturbed prompts for each response token is measured using a bounded-symmetric distance defined by Equation (1), which normalizes the difference to avoid scale bias.
How it works: Relative Stability Signal Construction
The token-level drifts are aggregated to create a relative stability signal within each prompt group. This is achieved by defining a group-wise standardization, where the mean and standard deviation of drift are computed over all valid response tokens in the group. The resulting score, denoted as 'u', measures within-group relative stability across different perturbation types. Crucially, only the negative part of this signal is retained: we retain only the negative part of the stability scores,
because stability alone does not imply correctness.
How it works: Semifactual Credit Augmentation and Policy Optimization
SCAPO modifies the standard GRPO advantage by adding a semifactual credit signal. The final token-level advantage is constructed as:
Aei,t = Ai + λ sgu−i,t, (Equation 9)
where 'Ai' is the original outcome-derived advantage from GRPO, and 'sgu−i,t' is the normalized stability score. This construction ensures that relatively unstable tokens receive lower advantages,
while stability alone does not grant additional credit. The augmentation strength and duration are controlled by parameters λ0 (correction strength) and N0 (augmentation duration), applied during an initial phase of training to shape trajectory selection before continuing with standard GRPO.
Key Findings and Contributions
The study demonstrates that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR.
On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improved AIME 2024–2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively, achieving the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among compared methods. Furthermore, the analysis shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights,
leading to a gain of 14.2 percentage points on Qwen3-4B-Base accuracy (from 15.8% to 30.0%) in d-filtered decoding, even when the final answer is correct. The ablation studies confirm that using negative-only corrections
yields the highest average accuracy across competition-level benchmarks, supporting the principle of selectively reducing credit for relatively unstable tokens without rewarding stability alone.
Computational Efficiency and Generalization
The probing and credit construction phases account for only 1.8% of the measured training runtime in the 4B experiment, as probes reuse sampled responses through teacher forcing. The results show that SCAPO achieves the highest accuracy on all three out-of-distribution benchmarks at both model scales, indicating that this signal is effective for improving generalization beyond the training domain. The method is designed to refine outcome rewards by refining token-level credit assignment based on semifactual sensitivity, proving that semifactual stability complements outcome rewards with an effective signal for finer-grained credit assignment in RLVR.
A LIMITATIONS
The training experiments are limited to mathematical reasoning with dense Qwen3 models of 1.7B and 4B parameters trained on DAPO-Math-17K. Future work is suggested to explore this credit augmentation in larger models, more diverse architectures, and domains beyond mathematical reasoning. The authors also note that the effectiveness of SCAPO at larger model and data scales remains to be evaluated.
AI USE STATEMENT
We used generative AI to generate semifactual perturbation data, polish writing, correct grammar, and assist with code debugging and script development.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems by implementing the concepts from Semifactual Credit-Augmented Policy Optimization (SCAPO), and what these improved systems could achieve:
The core improvements revolve around moving beyond coarse, response-level rewards (like those in Group Relative Policy Optimization - GRPO) to a finer, token-level credit assignment mechanism informed by semantic stability.
Specifically:
-
A new training objective that incorporates a
semifactual stability
signal during the early phases of Reinforcement Learning (RL). -
A mechanism to selectively reduce the advantage assigned to tokens that exhibit high probability drift when prompted with semifactual perturbations (i.e., prompts that preserve the core problem but alter irrelevant features).
The improved AI systems could achieve the following specific capabilities:
-
Improved reasoning accuracy on complex mathematical and quantitative problems by focusing training signals on tokens whose predictions are robust against minor, task-irrelevant prompt variations (like paraphrasing, typos, or scenario wrapping).
-
Enhanced generalization to out-of-distribution (OOD) benchmarks by training the model to recognize and rely only on
stable
reasoning pathways rather than spurious correlations present in the training distribution. -
More efficient policy optimization by using a small fraction of runtime for
semifactual probing
(measuring drift under perturbations) during an initial phase, rather than relying solely on expensive autoregressive rollout. -
Superior performance across diverse mathematical reasoning benchmarks (AIME, AMC, HMMT) and potentially other scientific domains due to the fine-grained credit assignment.
In summary, the improved AI system can move from simply predicting a final answer based on outcome rewards to learning a more robust internal reasoning process where it actively learns which intermediate steps (tokens) are reliable and which are likely artifacts of prompt sensitivity, leading to higher accuracy and better reliability in challenging mathematical tasks.
Sources
- Invariant Risk Minimization
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Soft Adaptive Policy Optimization
- Beyond Uniform Credit: Causal Credit Assignment for Policy Optimization
- Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR
- Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Qwen3 Technical Report
- Group Sequence Policy Optimization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks