Abstention vs. Hallucination: Benchmarking LLM Source Attribution for Scientific Citations
cs.CL, cs.AI, cs.IR
Submitted: 2024-05-03
Updated: 2026-09-15
Comments: accepted to 2026 13th International Conference on Data Science and Advanced Analytics (DSAA 2026)
Code: https://github.com/YashSaxena21/REASONS
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) increasingly generate citation-backed responses, yet citation hallucination remains a major challenge for trustworthy scientific information access.
Terminology
Abstract
Large language models (LLMs) increasingly generate citation-backed responses, yet citation hallucination remains a major challenge for trustworthy scientific information access. We introduce REASONS, a benchmark of 12,723 sentence-level citation instances spanning 12 arXiv subject categories, designed to evaluate scientific citation attribution under varying evidence conditions. We propose a dual-metric framework consisting of Abstention Rate (AR) and Hallucination Rate (HR) to characterize the trade-off between reliability and responsiveness. Using author-attribution and title-attribution tasks, we evaluate proprietary and open-source LLMs under zero-context, metadata-augmented, cascaded metadata-augmented prompting (CMP), retrieval-augmented, and adversarial settings. Advanced RAG lowers HR relative to Naive RAG (65.4% vs. 87.6%) but reduces AR from 5.0% to 0%. Under adversarial metadata, several systems exceed 85% HR, while retrieval-augmented variants frequently maintain near-zero abstention. Human evaluation of 1,000 outputs (κ=0.78) finds a 12.7:1 ratio of factual hallucinations to acceptable paraphrases. Our findings demonstrate that citation attribution systems should be evaluated not only for correctness but also for their ability to abstain appropriately under uncertainty. REASONS provides a benchmark and evaluation framework for studying attribution reliability in citation generation.
Sources
- Optimization-Inspired Learning with Architecture Augmentations and Control Mechanisms for Low-Level Vision
- Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network
- Evaluation of Attribution Bias in Generator-Aware Retrieval-Augmented Large Language Models
- On the Limitations of Large Language Models (LLMs): False Attribution
- CoTAR: Chain-of-Thought Attribution Reasoning with Multi-level Granularity
- Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models
- SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Facilitating Human-LLM Collaboration through Factuality Scores and Source Attributions
- Enabling Large Language Models to Generate Text with Citations
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Advancing Large Language Model Attribution through Self-Improving
- Atlas: Few-shot Learning with Retrieval Augmented Language Models
- Mistral 7B
- Active Retrieval Augmented Generation
- Source-Aware Training Enables Knowledge Attribution in Language Models
- Generalization through Memorization: Nearest Neighbor Language Models
- Evaluating Verifiability in Generative Search Engines
- S2ORC: The Semantic Scholar Open Research Corpus
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering