LLM Ghostbusters: Surgical Package Hallucination Suppression via Adaptive Unlearning
cs.CR, cs.AI, cs.CL, cs.LG
Submitted: 2026-05-01
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by/4.0/
The gist: Hallucinations remain an unsolved problem for LLMs, and package hallucinations are a particularly dangerous instance of this phenomenon.
Terminology
Abstract
Hallucinations remain an unsolved problem for LLMs, and package hallucinations are a particularly dangerous instance of this phenomenon. Package hallucinations occur during code generation when a model fabricates non-existent software packages, recommending imports and installation commands for fictional libraries. This creates a critical supply-chain vulnerability; an attacker can proactively register such packages on public registries with malicious payloads that are subsequently installed and executed by developers or autonomous agents. These hallucinations enable a class of package confusion attack known as slopsquatting. To address this issue, we present Adaptive Unlearning (AU), a post-deployment framework that surgically suppresses package hallucinations while preserving general model utility. AU introduces a hybrid token-level objective that simultaneously reinforces valid outputs and suppresses hallucinated ones. Combined with an adaptive discovery loop that continuously surfaces new hallucination-inducing contexts without human supervision, AU enables generalization to unseen prompts and hallucinations. We demonstrate that AU reduces package hallucination rates by 88%, while maintaining performance on standard coding benchmarks. Our analysis shows that distributional changes are concentrated on package-related generations, leaving general coding behavior largely unaffected and confirming that AU's effect is isolated to the targeted distribution. AU relies entirely on model-generated data and requires no human annotation, representing a post-deployment hallucination mitigation framework.
Sources
- Why Language Models Hallucinate
- Scaling Laws for Neural Language Models
- Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
- Learn to Forget: Machine Unlearning via Neuron Masking
- Program Synthesis with Large Language Models
- Constitutional AI: Harmlessness from AI Feedback
- Evaluating Large Language Models Trained on Code
- Forget-It-All: Multi-Concept Machine Unlearning via Concept-Aware Neuron Masking
- Strong Model Collapse
- Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- A Survey of Machine Unlearning
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
- ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
- Will we run out of data? Limits of LLM scaling based on human-generated data
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs