Attributable Post-Rationalization in RAG Citations: A Controlled Reproduction and an RLVR Comparison
cs.CL, cs.IR
Submitted: 2026-09-19
Updated: 2026-09-19
Comments: 11 pages, 1 figure, code available at: https://github.com/mehedikhan72/RAG-Post-Rationalization-RLVR-Comparison
Code: https://github.com/mehedikhan72/RAG-Post-Rationalization-RLVR-Comparison
License: http://creativecommons.org/licenses/by/4.0/
The gist: A RAG system can hand you the right answer and cite a source it did not actually use.
Terminology
Abstract
A RAG system can hand you the right answer and cite a source it did not actually use. Models output these unfaithful citations via post-rationalization: they write the answer first and then attach a citation to whatever passage looks close enough. Search agents are now trained with reinforcement learning from verifiable rewards (RLVR), which pays them for getting the answer right. We asked whether that training also teaches them to cite honestly. Improving an existing methodology with a required control, we compared an instruction-tuned model against three RLVR agents trained from it, on four question-answering datasets, using only free-tier Kaggle GPUs. Post-rationalization is everywhere: on Wikipedia-based questions roughly one citation in seven is unfaithful. RLVR does not fix it. The agents post-rationalize at their base model's rate, and one lands slightly worse. Rewarding correct answers buys nothing in citation faithfulness, so faithfulness has to be trained and measured on its own terms.
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering