FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision
cs.CL, cs.AI
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/verl-project/verl
License: http://creativecommons.org/licenses/by/4.0/
The gist: To reduce the hallucination risk caused by outcome-driven rewards in large language models trained through reinforcement learning with verifiable rewards, existing mitigation approaches introduce
Terminology
Abstract
To reduce the hallucination risk caused by outcome-driven rewards in large language models trained through reinforcement learning with verifiable rewards, existing mitigation approaches introduce process-level factual supervision. However, due to coarse-grained aggregation of factual signals and the lack of reliability assessment for these signals, they create a mismatch between fact verification and policy updates. We term this noisy factual credit assignment and decompose it into two aspects: credit localization ambiguity and credit reliability ambiguity. To address these issues, we propose FARCA (Fact-Aligned Reliability-Aware Credit Assignment), a policy optimization framework that transforms factual supervision into localized, reliability-weighted token-level training signals. FARCA achieves fine-grained credit localization by aligning the granularity of fact verification with that of policy updates. It further introduces counterfactual evidence attribution, which uses the dependence of a factual judgment on key evidence as an empirical proxy for verification reliability to compute reliability weights. These weights modulate factual rewards and local policy advantages, reducing the influence of potentially unreliable signals on policy optimization. Experiments across different models and multiple factual reasoning benchmarks show that FARCA significantly improves model factuality while preserving general reasoning capabilities.
Sources
- Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
- Learning to Reason for Factuality
- Reasoning Models Don't Always Say What They Think
- Evaluating Hallucinations in Chinese Large Language Models
- Training Verifiers to Solve Math Word Problems
- FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
- The Llama 3 Herd of Models
- FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Qwen2.5-Coder Technical Report
- GPT-4o System Card
- OpenAI o1 System Card
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
- Measuring short-form factuality in large language models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering