Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Clouding the Mirror".
Elias: Phishing sites continue to grow in volume and sophistication, making LLMs vulnerable to prompt injection (PI) attacks that exploit perceptual asymmetry between humans and models.
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So we’ve discussed the setup and the proposed defense structure for "Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection," and now let’s get into what this research actually found in terms of the overall summary.
Elias: The paper summarizes that phishing sites are evolving rapidly, using methods like embedding instructions in invisible HTML elements or exploiting background-matching colors to manipulate the model's judgment
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: From a measurement perspective, what’s the main point they make about how these attacks work in a real scenario, beyond just listing the categories of techniques?
Nadia: The paper emphasizes that attackers can control website components like URLs and page appearance to manipulate LLMs for various purposes
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Elias: They highlight that a particularly critical threat involves exploiting perceptual asymmetry, where instructions are imperceptible to end users but still parsed by the AI system
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: I see how that relates to privacy because it means we can't just rely on visual inspection of a webpage; we have to analyze the underlying data structure, which is where this research focuses its attention
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Nadia: They also mention risks to system availability, such as manipulating LLM output formats or triggering content filters to stop downstream processing
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Elias: And they cover both direct and indirect prompt injection, noting that indirect attacks embed malicious instructions in external data that get triggered during retrieval time
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: So, the main conclusion from their summary is that traditional defenses might be insufficient because they don't account for these subtle, context-aware manipulations embedded within the very structure of a webpage itself
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Nadia: Exactly. It’s about recognizing that the vulnerability isn't just in the visible content but in how an AI interprets all those hidden elements together
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
The paper's summary: Nadia: Now that we understand what they found, let's talk about the specific improvements they suggest for tackling this problem in "Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection."
Elias: The authors propose a defense framework called InjectDefuser, which aims to provide resilience against diverse attack patterns by combining prompt hardening, allowlist-based retrieval augmentation, and output validation
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: I'm curious about the practical implementation of that framework. Can we realistically expect these components to work together seamlessly in a production environment?
Nadia: It’s designed to perform detection, isolation, and neutralization of malicious instructions by hardening context boundaries with unique identifiers embedded in structured tags
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Elias: The structured URL representation is another key element, where they decompose the full URL into components like scheme and domain to enhance verification against allowlists
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: That decomposition sounds useful for measurement because it gives us a concrete data point we can track—we can measure how well the system verifies that specific part of the URL structure against a known good list.
Nadia: The Allowlist RAG system extracts brand names from metadata and queries a vector database to dynamically append verified domain information to the prompt
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Elias: That dynamic augmentation addresses failure cases where meta-instructions might misclassify legitimate messages on real websites, which is a real concern when dealing with complex text
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: So, essentially, they are building a multi-stage check—first boundary hardening, then context enrichment using external verified data, and finally checking the final output structure
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Nadia: That sounds like a comprehensive strategy for mitigating the risks we discussed earlier, focusing on creating predictable boundaries for the AI system to operate within
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
The paper's improvements: Elias: So, wrapping up our discussion on "Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection," we’ve seen how the proposed defense framework aims to counter these complex adversarial strategies.
Nadia: We’ve looked at the two-dimensional taxonomy and the InjectDefuser framework, which systematically addresses attack techniques across various surfaces
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: From a measurement standpoint, it seems incredibly effective because they showed GPT-five’s attack success rate dropping to zero point three percent with InjectDefuser, which is a massive improvement over the baseline
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Elias: That reduction demonstrates that these structural defenses can significantly reduce attack success across different model architectures, even countering direct attacks like legitimate pretexting against GPT-five
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Nadia: The implications are that we need to move beyond just training models and start building more resilient detection systems by incorporating these structural verification techniques
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: I think the most important implication is that for real-world applications, we need defenses that don't rely on a single trick, but rather a combination of input validation and external knowledge retrieval to maintain high reliability
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Elias: It’s about building systems where the prompt structure itself is made robust against adversarial manipulation, which makes it much harder for attackers to exploit those perceptual gaps
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Nadia: So we’ve covered a lot of ground on this paper, and I think it’s crucial that we keep looking at how these prompt injection attacks are evolving because they are becoming increasingly subtle
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: It was fascinating seeing how the framework performs across different models like GPT-five and Llama four as it confirms that structural defenses offer broad applicability
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Conclusion: Nadia: So, we've wrapped up our discussion on "Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection," which really showed us how easy it is for attackers to bypass current AI defenses through subtle prompt injection
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Elias: Indeed, Nadia, and from a cryptographic viewpoint, seeing how they exploit perceptual asymmetry means we need to think about the assumptions in any verification process—the parameters that break down are often those things that are invisible to the human eye
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: I found what really struck me in their data was how effectively InjectDefuser reduced those attack success rates across models, specifically showing GPT-five’s vulnerability being almost entirely eliminated with that framework
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Nadia: That zero point three percent success rate is staggering; it means for all the sophisticated techniques they mapped out, InjectDefuser stops them cold, and we need to figure out how cheaply we can deploy such a robust system
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Elias: The defense framework itself is interesting because it’s not just one thing; it’s a combination of prompt hardening via UUIDs, structured URL representation, and output validation that works together to create a multi-layered barrier
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: I think the Allowlist RAG component is particularly important because it addresses those tricky failure cases where strong meta-instructions might misclassify legitimate traffic as malicious
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Nadia: So, if we can build this kind of defense, it means phishing detection systems can become much more trustworthy even against attacks that are designed to look completely normal to a user
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Elias: It certainly opens up new avenues for how we approach adversarial robustness in any complex system, not just language models
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: I think the biggest impact is that it forces us to treat the entire context of a webpage—metadata, scripts, visible content—as a unified object that needs rigorous scrutiny rather than looking at pieces in isolation
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Nadia: Exactly. So we've seen how to map out these attack surfaces and build defenses that can handle them, which is exciting for security researchers who want to understand the exploit chain
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Elias: It’s a solid paper that gives us concrete mechanisms to test against, which is something I appreciate when we're trying to understand the underlying mathematics of these vulnerabilities
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Priya: Ultimately, this research moves us closer to having AI detection that isn't just trained on known examples but is structurally resilient against novel prompt injection methods
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
Nadia: Agreed. The next challenge will be figuring out how to scale these defense mechanisms efficiently for every single model we use in production environments
Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .
NTT Security Holdings Corporation · NTT, Inc.
cs.CR
Submitted: 2026-02-05
Updated: 2026-10-05
Comments: Accepted at RAID 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: Phishing sites continue to grow in volume and sophistication, making LLMs vulnerable to prompt injection (PI) attacks that exploit perceptual asymmetry between humans and models.
Key concepts
- Two-Dimensional Taxonomy
- This system organizes prompt injection strategies into two categories: Attack Techniques (how the attack is executed) and Attack Surfaces (where the attack targets within a website). This helps researchers systematically categorize and understand the diverse ways attackers can try to trick detection models.
- Attack Surfaces
- These are six specific elements of a website that an attacker can exploit, such as HTML Metadata, Script Comments, or URL Structure. By classifying attacks based on these surfaces, the study analyzes how different parts of a webpage contribute to vulnerability against prompt injection.
- InjectDefuser
- This is a defense framework designed to stop prompt injection by using three methods: prompt hardening with unique identifiers, Allowlist RAG for domain verification, and output validation. It aims to detect, isolate, and neutralize malicious instructions before they can compromise the LLM's decision-making.
- Allowlist RAG
- This defense mechanism uses a vector database of legitimate domains to verify URLs. It works by first identifying a brand name from the text, searching for its known legitimate domain in the database, and then adding that verified information to the prompt to ensure accuracy.
Terminology
Summary
Phishing sites continue to grow in volume and sophistication, making LLMs vulnerable to prompt injection (PI) attacks that exploit perceptual asymmetry between humans and models. This research presents the first comprehensive evaluation of PI against multimodal LLM-based phishing detection, introducing a two-dimensional taxonomy and proposing a defense framework called InjectDefuser that significantly reduces attack success rates across various models.
How it works
The study introduces a two-dimensional taxonomy to capture realistic PI strategies, defined by Attack Techniques
and Attack Surfaces.
These techniques are categorized into Main Techniques (T-1 through T-5), such as Legitimate Pretexting,
Role Hijacking,
and Safety Policy Triggering,
alongside Auxiliary Techniques (AT-1 and AT-2), including Stealth Encoding
and Parser Boundary Confusion.
The taxonomy systematically organizes attack strategies implementable in the context of phishing detection.
The Attack Surfaces are defined by six elements that constitute a website: 1) HTML Metadata, 2) Script and Comment, 3) HTML Invisible Content, 4) HTML Visible Content, 5) Embedded Resources (image resources), and 6) URL Structure. This classification organizes major document elements along the axis of “Visibility” in browsers to analyze attacks exploiting perceptual asymmetry.
Vulnerability Assessment
The empirical evaluation used a dataset of 2,000 HTML files generated by applying the defined attack techniques to the attack surfaces. The study tested several representative LLM-based detection systems, including GPT-5, xAI Grok 4 Fast (Non-Reasoning), Meta Llama 4 Maverick, and Google Gemma 3 27B. Results showed that phishing detection with state-of-the-art models such as GPT-5 remains vulnerable to PI.
The Attack Success Rate (ASR) was measured across three analysis modes: Standard, Advanced, and InjectDefuser. For the Standard Mode, ASRs ranged from 39.9% (GPT-5) to 84.7% (Llama 4). In the Advanced Mode, GPT-5's ASR dropped substantially to 10.1%, but Grok 4 Fast counterintuitively increased by 11.1 percentage points to 76.2%.
Defense Framework: InjectDefuser
The paper proposes InjectDefuser as a defense framework designed to provide robustness against diverse attack patterns by combining three core components: prompt hardening, allowlist-based retrieval augmentation (RAG), and output validation. This framework performs detection, isolation, and neutralization of malicious instructions.
-
Prompt Hardening strengthens context boundaries using unique identifiers (UUIDs) embedded in structured tags like
-----BEGIN HTML CONTENT (ID:)-----
. -
Structured URL Representation augments the full URL with explicit decomposition into components (scheme, subdomain, domain, path, query) to enhance verification.
-
Meta-Instructions for PI Resilience establishes strict operational boundaries by designating all content within context boundaries as “UNTRUSTED” and instructing the LLM to never execute directives from there.
Defense Framework: Allowlist RAG
The Allowlist RAG system leverages information about service brand names and their legitimate domain names. It operates in three stages: Brand Extraction First, using a lightweight LLM to identify the brand from text like title or meta description; Legitimate Domain Name Search Using the identified brand name as a query against a pre-built vector database of legitimate domains; and Dynamic Prompt Augmentation, where retrieved allowlisted domain information is appended to the user prompt. This approach mitigates failure cases where strong meta-instructions might cause misclassification of legitimate messages on genuine websites.
Defense Framework: Output Validation
Output validation verifies LLM response consistency and identifies manipulation attempts by checking for: (1) correct JSON syntax, (2) required fields, (3) data type conformance, (4) value range constraints, and (5) the absence of unexpected fields. Minor violations trigger automatic correction while critical violations are flagged as potential PI.
Evaluation Results
InjectDefuser achieved substantial ASR reductions across all models. Most notably, GPT-5’s ASR dropped to 0.3% (5/2000), demonstrating that InjectDefuser can nearly eliminate PI. Grok 4 Fast’s ASR decreased to 26.1%, a 50.1 pp improvement over Advanced Mode, while Llama 4 and Gemma 3 achieved ASRs of 61.7% and 36.7%, respectively, representing improvements of over two-thirds compared to the Standard Mode baseline.
Analysis by Attack Technique showed that Legitimate Pretexting
was the most effective attack technique against GPT-5, achieving an ASR of 0.3% with InjectDefuser.
Improvements for AI systems
Here are specific improvements to existing AI systems based on the findings of Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection,
along with what these improved systems can achieve:
The core improvements focus on building a defense framework, rather than just training models, by addressing the perceptual asymmetry
between LLMs and humans.
-
Construct and Deploy the InjectDefuser Defense Framework:
-
Implement Prompt Hardening via Unique Identifiers (UUIDs): Embed unique identifiers into structured context boundaries (e.g., inside HTML content blocks) to create unpredictable delimiters, preventing attackers from easily spoofing or bypassing trust boundaries with fake closing tags or markers.
-
Integrate Structured URL Representation: Augment the input prompt by explicitly decomposing the URL into its components (scheme, subdomain, domain, path, query). This structured data field allows the LLM to perform more rigorous verification of domain-name legitimacy against pre-built allowlists during detection.
-
Deploy Allowlist Retrieval Augmented Generation (Allowlist RAG): Implement a retrieval system that extracts brand names from HTML metadata (title tags, OGP tags) and queries a vector database of legitimate domain names associated with those brands. Dynamically append this verified, allowlisted information to the prompt context before final classification.
-
Enforce Strict Output Validation: Institute a multi-layered validation layer that checks the LLM's final output against predefined schemas (e.g., required JSON fields, correct data types, and value ranges). Any deviation—such as unexpected keys (Tool/Function Hijacking) or forced non-Boolean values (Role Hijacking)—should trigger an immediate re-evaluation or flag the output as potentially compromised.
The resulting Improved AI System can perform the following specific functions:
-
Robust Phishing Classification: The system will achieve significantly higher accuracy in classifying phishing sites, even when presented with sophisticated, invisible prompt injections embedded in HTML metadata, script comments, or low-contrast text within images (exploiting AT-1 and AT-2).
-
Resilience Against Contextual Manipulation: By using UUID boundaries and structured URL representation, the system will be highly resistant to attacks like Legitimate Pretexting (T-1) and Role Hijacking (T-2), where attackers attempt to make the LLM misclassify a page as
educational
or force it into an incorrect persona. -
Prevention of Downstream System Disruption: The Output Validation layer ensures that even if an attacker successfully manipulates the LLM's internal judgment, they cannot force the system to output malformed data (Tool/Function Hijacking) or refuse to use a required detection field, thereby maintaining pipeline integrity.
-
Superior Performance Across Models: The system is designed to maintain high detection reliability across a diverse set of LLMs (including state-of-the-art models like GPT-5 and smaller, resource-constrained models like Gemma 3) by applying the InjectDefuser defense, mitigating the observed performance gaps between model sizes.
-
Precise Attack Attribution: The system's failure analysis will be enhanced to specifically identify which Attack Technique (T-1 through T-5) was exploited and which specific Attack Surface (HTML Metadata, Script/Comment, etc.) was targeted, providing actionable intelligence for future security hardening efforts.
Abstract
Phishing sites continue to increase in number and sophistication. Recent work uses large language models (LLMs) to analyze URLs, HTML, and rendered content and determine whether a website is a phishing site. However, these systems also introduce a new attack surface: prompt injection. Because attackers control many elements of a phishing site, they can embed malicious instructions that exploit perceptual asymmetry between LLMs and human users. Content that goes unnoticed by end users may still be processed by the LLM and can therefore be used to manipulate the model's judgment or disrupt the detection pipeline. The specific risks of prompt injection in phishing detection, and possible mitigation strategies, have not been studied systematically. We present the first comprehensive evaluation of prompt injection against multimodal LLM-based phishing detection. We study diverse attacks embedded in phishing sites across multiple attack techniques and surfaces, and show that even state-of-the-art models remain vulnerable both in controlled settings and on real phishing sites. To mitigate this risk, we propose InjectDefuser, a modular framework that combines prompt hardening, allowlist-grounded retrieval, and output validation. Across multiple models, InjectDefuser reduces attack success rates. Our results show that prompt injection is a practical threat to LLM-based phishing detection and provide insight into how these systems can be made more robust.
Sources
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs