Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
cs.AI
Submitted: 2025-09-27
Updated: 2026-09-18
Comments: 15 pages, 3 figures, 4 tables. Code and dataset available at https://github.com/ayushgupta4897/FGA
Code: https://github.com/ayushgupta4897/FGA
License: http://creativecommons.org/licenses/by/4.0/
The gist: "The greatest enemy of knowledge is not ignorance, it is the illusion of knowledge." Large Language Models have conquered natural language but remain prisoners of their own probabilistic
Terminology
Abstract
"The greatest enemy of knowledge is not ignorance, it is the illusion of knowledge." Large Language Models have conquered natural language but remain prisoners of their own probabilistic nature--confidently hallucinating facts they never truly knew. We present Fact Grounded Attention (FGA), a novel architectural modification that transforms unreliable language models into deterministic truth tellers by injecting verifiable knowledge directly into the attention mechanism. Unlike existing approaches that patch hallucinations after generation or prepend retrieved text, FGA intervenes at the mathematical heart of the transformer--the pre-softmax attention scores--creating a model that cannot hallucinate when facts exist in its knowledge base. Our experiments across 1,107 technical queries spanning smartphones, laptops, and electric vehicles demonstrate a transformation from 6.3% accuracy in vanilla Llama 3.2 to 99.7% accuracy with FGA. More critically, knowledge updates occur in under one second without retraining, compared to hours for parameter editing approaches. FGA doesn't just reduce hallucination--it eliminates it entirely for verifiable facts, marking a fundamental shift from probabilistic approximation to deterministic precision in neural language generation.
Sources
- Deep Patch Visual Odometry
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- LLaMA: Open and Efficient Foundation Language Models
- Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection