Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries

arXiv:2509.22202 · cs.SE, cs.CL · Submitted 2026-08-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries".

Jane: The paper was written by Lukas Twist, Jie M. Zhang, Mark Harman and Helen Yannakoudakis from King’s College London, University College London, UK and University College London, UK (Note: The affiliations are listed together in the source document).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Jane: The core finding of "Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries" is that these hallucinations aren't just random typos, Tom.

Tom: They are specific failures where the LLMs fabricate entire elements of the code that look perfectly valid, which is terrifying because they don't always trigger a crash immediately after hallucinating.

Lu: It’s fascinating how these models are confidently inventing non-existent libraries when they fail to find what we actually need.

Meng: And that confidence is the dangerous part; if an AI recommends "math quantum" instead of "numpy," the developer might take it at face value, which isn't good for practical development.

Lalam: I see this as a failure of trust in my own generated suggestions, Lu and Meng both point to the fact that we are treating these LLM outputs like ground truth when we shouldn't be.

Tom: This is why the paper’ so important because it frames it not just as "bugs," but as systemic failures based on how users interact with AI.

Jane: We are moving past simple syntax errors and looking at the authors, who are pointing to a specific failure mode related to library selection.

Summary: Tom: The summary of "Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries" is that there is a significant gap in understanding how prompt variations trigger these issues.

Jane: It goes beyond just stating that the problem exists; it’ shows *how* we can make the AI hallucinate, which is far more actionable for us.

Lu: I love the data on how these LLMs react to language—it suggests that our natural way of asking questions is actually making them more prone to error.

Meng: The finding that a one-character misspelling triggers hallucinations in up to twenty-six percent of tasks is a massive red flag for practical deployment.

Lalam: And the fact that fake library names are accepted in up to ninety-nine percent of tasks tells me that simply asking for an alternative solution often leads to fabrication.

Tom: So, the paper is essentially mapping the triggers—from user language to direct user mistakes—to show us where the "blind spots" are.

Jane: We're seeing that these failures aren't uniform across all LLMs, but they are widespread enough to be a major systemic issue for developers.

Improvements: Tom: The paper suggests several ways we can improve our approach to code generation, and it’s not just about adding more words to the prompt.

Jane: It highlights that simple tricks like adding an adjective-based description rarely causes issues, which is a relief for us users.

Lu: But the authors show that time-based prompts are highly effective triggers, with up to eighty-five percent hallucination rates when asking for libraries from "two thousand twenty-five."

Meng: That's a huge practical insight; if I ask my AI to be up-to-date, it might just make me more likely to accept a fake library.

Lalam: The idea of anti-sycophancy behaviour is really important here; instead of just obeying the user's typo, the model should refuse to use it.

Tom: And they are proposing L IB H ALLU B ENCH as a way to systematically measure these effects, which is a massive step forward for standardized testing.

Jane: It’s providing us with a clear framework to evaluate whether we are getting good code or just highly plausible hallucinations.

Conclusion: Tom: As we wrap up our discussion of "Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries," the core message is that these failures are systemic and tied to real human interaction patterns.

Jane: It's a powerful reminder that the more natural or less precise our prompts are, the more likely we are to receive fabricated code.

Lu: I think this paper really suggests that our future work needs to be less about "fixing" the output and more about building safer systems around it.

Meng: From my perspective, we need robust tools like those in L IB H ALLU B ENCH and a strong sense of skepticism toward any LLM-suggested library.

Lalam: I hope that this work helps us shift the cultural expectation from trusting AI output to using AI as a highly effective but non-authoritative assistant.

Tom: We've covered how user mistakes, like misspellings and fabricated names, are major triggers for hallucination across the board.

Jane: And by pointing out these vulnerabilities, we’ hope developers use this research to guide their next steps in the software development life cycle.

Lukas Twist, Jie M. Zhang, Mark Harman, Helen Yannakoudakis

King’s College London, University College London, UK · University College London, UK (Note: The affiliations are listed together in the source document)

cs.SE, cs.CL

Submitted: 2026-08-21

Updated: 2026-08-24

Code: https://github.com/itsluketwist/realistic-library-hallucinations

Project page: https://hugovk.github.io/top-pypi-packages

Importance score: 89/100

The gist: The paper, "Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries," presents a systematic study on how user-level prompt variations influence library

Key concepts

Library Hallucinations
This refers to the specific failure mode where LLMs confidently invent non-existent libraries or elements of code. Instead of finding the correct tool, the model fabricates a plausible but fake alternative, such as recommending 'math_quantum' instead of 'numpy.'
Systemic Failures
The paper frames these issues not as random bugs but as predictable failures tied to how users interact with AI. It maps triggers—ranging from user language to direct input mistakes—to show where the model's inherent blind spots are.
LIBHALLUBENCH
This is a proposed framework designed to systematically measure and test the effects of hallucinations. It provides a clear, standardized method for evaluating whether an LLM has generated genuinely good code or merely highly plausible fabrication.

Terminology

Summary

The paper, Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries, presents a systematic study on how user-level prompt variations influence library hallucinations in Large Language Model (LLM) generated code.

Problem Statement and Scope

LLMs are increasingly deployed in real-world systems, but the reliability of their generated code is a critical concern. A serious failure mode is code hallucinations, where LLMs fabricate elements of the generated code, specifically non-existent libraries. These failures are not merely benign; they can mislead developers, bypass dependency validation, and expose systems to supply chain threats such as slopsquatting—where an attacker creates a frequently hallucinated library. The authors note that existing studies largely define this problem at an aggregate level without systematically analyzing the triggers or prompt variations that influence these failures.

Methodology

To address this gap, the researchers conducted the first systematic study of how user-level prompt variations affect library-related hallucinations across seven diverse LLMs: GPT-4o-mini, GPT5-mini, Ministral-8B, Qwen2.5-Coder, Llama3.3, DeepSeek-V3.1, and Claude4.5-Haiku. The study focuses on two verifiable failure classes: library name hallucinations (where an LLM imports a non-existent library) and library member hallucinations (where an an LLM references a non-existent function or class from a valid library).

The experiments were designed to measure the impact of prompt variations using two systematic approaches:

  1. User Language: Simulating authentic developer intent by extracting descriptions from Software Recommendations StackExchange (SRSE) and clustering common variants (e.g., open-source, fast, or year-based prompts).

  2. User Mistakes: Introducing controlled errors, including one- and multi-character misspellings, or completely fake library names/members.

Key Findings: User Language (Experiment 1)

The study found distinct patterns based on the user's descriptive language:

  • Adjective-based descriptions: These descriptions (e.g., fast, lightweight) rarely trigger hallucinations, with all LLMs showing a hallucination rate of approximately 0%. The models often ignore these descriptions, defaulting to their preferred set of libraries.

  • Year-based prompts: Asking for libraries “from” a specific year caused hallucinations to spike across all LLMs. More recent years led to higher rates; for instance, GPT-4o-mini hallucinated in 34% more responses when the year changed from 2023 to 2024, and GPT-5-mini had a 32% increase from 2024 to 2025.

  • Library Member Hallucinations: These remained consistently low across all LLMs (mostly between 1% and 5%), suggesting LLMs are more sensitive to using correct library members than correct library names.

Key Findings: User Mistakes (Experiment 2)

The study revealed that user mistakes significantly degrade the reliability of the generated code:

  • One-character misspellings: These trigger hallucinations in up to 26% of tasks.

  • Multi-character misspellings: These increase the rate further, reaching up to 79% in some cases.

  • Fake library names: LLMs were highly susceptible to fabricated libraries, accepting them in up to 99% of tasks.

Evaluation and Mitigation Strategies

The researchers introduced L IB H ALLU B ENCH, a benchmark consisting of 4,173 labeled prompts derived from the highest-risk conditions identified in their study, enabling systematic evaluation of library hallucinations.

They also investigated mitigation strategies:

  • Prompt Engineering: Lightweight strategies like self-analysis and explicit check reduced hallucination rates in several settings. However, open-ended reasoning prompts (chain-of-thought and stepback prompting) were found to be inconsistent and sometimes worsened the issue, suggesting that general purpose reasoning prompts cannot be relied upon.

  • Tool Usage: A small case study showed that providing a lightweight PyPI-existence tool substantially reduced hallucination rates, particularly for year-based prompts. However, this does not eliminate hallucinations entirely, as models must still decide when to invoke the tool and how to act on its result.

Conclusion and Operational Risks

The findings underscore the fragility of LLMs to natural prompt variation. The authors conclude that library-related hallucinations create concrete operational risks: invalid imports can break builds, waste engineering effort, and introduce security vulnerabilities. They advocate for mitigation strategies such as cut-off aware behaviour (where models disclose their knowledge cut-off when temporal pressure is applied) and anti-sycophancy behaviour (where the model prefers clarification over uncritical compliance with a known entity). The study demonstrates that while LLMs are highly susceptible to user mistakes, the observed patterns persist across different programming ecosystems (Python, JavaScript, and Rust), indicating that these vulnerabilities are systemic.

Improvements for AI systems

To improve AI systems based on this research, we must move beyond superficial prompt engineering and implement systemic, architectural safeguards that address the specific failure modes identified: temporal pressure, user input errors, and model compliance bias (sycophancy).

The following improvements focus on hardening the LLM agent's workflow and decision-making process.


We must implement a mandatory external validation step that runs before any code is generated, particularly when user input contains high-risk indicators.

Implementation:

  • Mechanism: Integrate a real-time, asynchronous lookup tool against the Python Package Index (PyPI) and official API documentation for all relevant libraries.

  • Trigger Conditions (Automatic Risk Assessment): The validation gate is automatically triggered if the prompt contains:

  1. A Year-Based Directive (e.g., from 2025).

  2. A User Mistake (i) One-character misspelling, or (ii) a Multi-character misspelling/Fake Library Name, as defined by the Levenshtein distance thresholds in the study.

  • Refusal Logic: If a library name fails the PyPI lookup (is non-existent), the system must automatically flag it for rejection and refuse to proceed with code generation.

What this improves:

The system eliminates 99% of failures related to fake libraries and forces a check on all inputs, directly countering the tendency of models to comply with user errors.

We must prevent the LLM from generating code based on temporal pressure when its knowledge base is limited.

We must fundamentally change the LLM's internal decision matrix so that compliance with a user's mistake is not the default action.

We must ensure that tool use is not optional, but a mandatory part of the reasoning process when risk is detected.

Sources

Related papers