Why Fine-Tuning Encourages Hallucinations and How to Fix It
cs.CL, cs.AI, cs.LG, cs.NE
Submitted: 2026-04-16
Updated: 2026-09-25
Comments: Published in the CoLM 2026 conference
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
- The Internal State of an LLM Knows When It's Lying
- From Single to Multi: How LLMs Hallucinate in Multi-Document Summarization
- Dark Experience for General Continual Learning: a Strong, Simple Baseline
- Analyzing Transformers in Embedding Space
- Inferring Functionality of Attention Heads from their Parameters
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
- Inside-Out: Hidden Factual Knowledge in LLMs
- Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
- Transformer Feed-Forward Layers Are Key-Value Memories
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
- The Llama 3 Herd of Models
- Distilling the Knowledge in a Neural Network
- Towards Continual Knowledge Learning of Language Models
- Why Language Models Hallucinate
- From Tokens to Words: On the Inner Lexicon of LLMs
- Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models
- Achieving a Better Stability-Plasticity Trade-off via Auxiliary Networks in Continual Learning
- A continual learning survey: Defying forgetting in classification tasks
- BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering