Small Updates, Big Doubts: Does Parameter-Efficient Fine-tuning Enhance Hallucination Detection ?
cs.CL, cs.AI
Submitted: 2026-01-17
Updated: 2026-08-29
Comments: 18 pages, 13 figures, 8 tables
Journal ref: Transactions on Machine Learning Research (TMLR), 08/2026 Transactions on Machine Learning Research
Code: https://github.com/google-research-datasets/natural-questions
License: http://creativecommons.org/licenses/by/4.0/
The gist: Parameter-efficient fine-tuning (PEFT) methods are widely used to adapt large language models (LLMs) to downstream tasks and are often assumed to improve factual correctness.
Terminology
Abstract
Parameter-efficient fine-tuning (PEFT) methods are widely used to adapt large language models (LLMs) to downstream tasks and are often assumed to improve factual correctness. However, how the parameter-efficient fine-tuning methods affect hallucination behavior remains insufficiently understood, especially on QA datasets. In this work, we systematically investigate the impact of PEFT on hallucination detection through a comprehensive empirical study across three open-weight LLM backbones and three fact-seeking QA benchmarks. For each model, we evaluate performance using seven unsupervised hallucination detection methods spanning three complementary approaches: semantic consistency based detectors, confidence based detectors, and entropy based detectors. This multifaceted evaluation enables us to characterize how PEFT reshapes uncertainty across different detection paradigms. In conclusion, our experimental results show that PEFT consistently strengthens hallucination detection ability, substantially improving AUROC across a wide range of hallucination detectors. Besides, further analyses using linear probes and representation diagnostics indicate that PEFT methods primarily reshapes how uncertainty is encoded and surfaced, comparing with injecting new factual knowledge into the models.
Sources
- The Llama 3 Herd of Models
- Mistral 7B
- Language Models (Mostly) Know What They Know
- Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
- LM-Polygraph: Uncertainty Estimation for Language Models
- Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
- DoRA: Weight-Decomposed Low-Rank Adaptation
- Qwen3 Technical Report
- Are Reasoning Models More Prone to Hallucination?
- PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering