Probing for Knowledge Attribution in Large Language Models
cs.CL, cs.AI
Submitted: 2026-02-26
Updated: 2026-09-10
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, where the model misuses provided context, and factuality
Terminology
Abstract
Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, where the model misuses provided context, and factuality violations, where answers reflect errors in internal knowledge. Proper mitigation depends on knowing which source drives each answer. We study contributive attribution, i.e. the classification of the dominant knowledge source behind each output, and show that a simple linear probe trained on hidden representations can reliably identify it. We introduce AttriWiki, a self-supervised pipeline that automatically generates labelled training data by prompting models to recall withheld entities from memory or read them from context without relying on knowledge conflicts. Probes trained on AttriWiki achieve up to 0.96 Macro- F 1 on Llama-3.1-8B, Mistral-7B, and Qwen-7B, transfer to SQuAD and WebQuestions with 0.94-0.99 Macro- F 1, and generalise zero-shot to Tighidet et al. (2024)'s benchmark, outperforming their probe on conflicting settings without retraining. Furthermore, attribution mismatches raise error rates by up to 70%, though correct attribution does not guarantee correct answers, pointing to the need for broader detection frameworks.
Sources
- On Mechanistic Circuits for Extractive Question-Answering
- Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models
- Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models
- The Llama 3 Herd of Models
- Chainpoll: A high efficacy method for LLM hallucination detection
- Mistral 7B
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Source-Aware Training Enables Knowledge Attribution in Language Models
- A Survey of Automatic Hallucination Evaluation on Natural Language Generation
- Probing Language Models on Their Knowledge Source
- Qwen2.5 Technical Report
- Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations
- Functional Abstraction of Knowledge Recall in Large Language Models
- Lynx: An Open Source Hallucination Evaluation Model
- Knowledge Conflicts for LLMs: A Survey
- A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
- Evaluating the External and Parametric Knowledge Fusion of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering