Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection

summary

Video file (mp4)

The gist

Phishing sites continue to grow in volume and sophistication, making LLMs vulnerable to prompt injection (PI) attacks that exploit perceptual asymmetry between humans and models.

In short

This research evaluated prompt injection (PI) attacks against multimodal LLMs used for phishing detection. The study developed a two-dimensional taxonomy to classify attack techniques and surfaces, testing several models like GPT-5 and Llama 4. A new defense framework, InjectDefuser, combining prompt hardening, allowlist RAG, and output validation significantly reduced attack success rates across all tested models.

Key concepts

Two-Dimensional Taxonomy
This system organizes prompt injection strategies into two categories: Attack Techniques (how the attack is executed) and Attack Surfaces (where the attack targets within a website). This helps researchers systematically categorize and understand the diverse ways attackers can try to trick detection models.
Attack Surfaces
These are six specific elements of a website that an attacker can exploit, such as HTML Metadata, Script Comments, or URL Structure. By classifying attacks based on these surfaces, the study analyzes how different parts of a webpage contribute to vulnerability against prompt injection.
InjectDefuser
This is a defense framework designed to stop prompt injection by using three methods: prompt hardening with unique identifiers, Allowlist RAG for domain verification, and output validation. It aims to detect, isolate, and neutralize malicious instructions before they can compromise the LLM's decision-making.
Allowlist RAG
This defense mechanism uses a vector database of legitimate domains to verify URLs. It works by first identifying a brand name from the text, searching for its known legitimate domain in the database, and then adding that verified information to the prompt to ensure accuracy.

Terminology used across episodes

This episode discusses

The paper

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection · Read on arXiv

NTT Security Holdings Corporation · NTT, Inc.

Phishing sites continue to increase in number and sophistication. Recent work uses large language models (LLMs) to analyze URLs, HTML, and rendered content and determine whether a website is a phishing site. However, these systems also introduce a new attack surface: prompt injection. Because attackers control many elements of a phishing site, they can embed malicious instructions that exploit perceptual asymmetry between LLMs and human users. Content that goes unnoticed by end users may still be processed by the LLM and can therefore be used to manipulate the model's judgment or disrupt the detection pipeline. The specific risks of prompt injection in phishing detection, and possible mitigation strategies, have not been studied systematically. We present the first comprehensive evaluation of prompt injection against multimodal LLM-based phishing detection. We study diverse attacks embedded in phishing sites across multiple attack techniques and surfaces, and show that even state-of-the-art models remain vulnerable both in controlled settings and on real phishing sites. To mitigate this risk, we propose InjectDefuser, a modular framework that combines prompt hardening, allowlist-grounded retrieval, and output validation. Across multiple models, InjectDefuser reduces attack success rates. Our results show that prompt injection is a practical threat to LLM-based phishing detection and provide insight into how these systems can be made more robust.

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Clouding the Mirror".

Elias: Phishing sites continue to grow in volume and sophistication, making LLMs vulnerable to prompt injection (PI) attacks that exploit perceptual asymmetry between humans and models.

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So we’ve discussed the setup and the proposed defense structure for "Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection," and now let’s get into what this research actually found in terms of the overall summary.

Elias: The paper summarizes that phishing sites are evolving rapidly, using methods like embedding instructions in invisible HTML elements or exploiting background-matching colors to manipulate the model's judgment

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: From a measurement perspective, what’s the main point they make about how these attacks work in a real scenario, beyond just listing the categories of techniques?

Nadia: The paper emphasizes that attackers can control website components like URLs and page appearance to manipulate LLMs for various purposes

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Elias: They highlight that a particularly critical threat involves exploiting perceptual asymmetry, where instructions are imperceptible to end users but still parsed by the AI system

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: I see how that relates to privacy because it means we can't just rely on visual inspection of a webpage; we have to analyze the underlying data structure, which is where this research focuses its attention

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Nadia: They also mention risks to system availability, such as manipulating LLM output formats or triggering content filters to stop downstream processing

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Elias: And they cover both direct and indirect prompt injection, noting that indirect attacks embed malicious instructions in external data that get triggered during retrieval time

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: So, the main conclusion from their summary is that traditional defenses might be insufficient because they don't account for these subtle, context-aware manipulations embedded within the very structure of a webpage itself

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Nadia: Exactly. It’s about recognizing that the vulnerability isn't just in the visible content but in how an AI interprets all those hidden elements together

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

The paper's summary: Nadia: Now that we understand what they found, let's talk about the specific improvements they suggest for tackling this problem in "Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection."

Elias: The authors propose a defense framework called InjectDefuser, which aims to provide resilience against diverse attack patterns by combining prompt hardening, allowlist-based retrieval augmentation, and output validation

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: I'm curious about the practical implementation of that framework. Can we realistically expect these components to work together seamlessly in a production environment?

Nadia: It’s designed to perform detection, isolation, and neutralization of malicious instructions by hardening context boundaries with unique identifiers embedded in structured tags

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Elias: The structured URL representation is another key element, where they decompose the full URL into components like scheme and domain to enhance verification against allowlists

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: That decomposition sounds useful for measurement because it gives us a concrete data point we can track—we can measure how well the system verifies that specific part of the URL structure against a known good list.

Nadia: The Allowlist RAG system extracts brand names from metadata and queries a vector database to dynamically append verified domain information to the prompt

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Elias: That dynamic augmentation addresses failure cases where meta-instructions might misclassify legitimate messages on real websites, which is a real concern when dealing with complex text

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: So, essentially, they are building a multi-stage check—first boundary hardening, then context enrichment using external verified data, and finally checking the final output structure

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Nadia: That sounds like a comprehensive strategy for mitigating the risks we discussed earlier, focusing on creating predictable boundaries for the AI system to operate within

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

The paper's improvements: Elias: So, wrapping up our discussion on "Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection," we’ve seen how the proposed defense framework aims to counter these complex adversarial strategies.

Nadia: We’ve looked at the two-dimensional taxonomy and the InjectDefuser framework, which systematically addresses attack techniques across various surfaces

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: From a measurement standpoint, it seems incredibly effective because they showed GPT-five’s attack success rate dropping to zero point three percent with InjectDefuser, which is a massive improvement over the baseline

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Elias: That reduction demonstrates that these structural defenses can significantly reduce attack success across different model architectures, even countering direct attacks like legitimate pretexting against GPT-five

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Nadia: The implications are that we need to move beyond just training models and start building more resilient detection systems by incorporating these structural verification techniques

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: I think the most important implication is that for real-world applications, we need defenses that don't rely on a single trick, but rather a combination of input validation and external knowledge retrieval to maintain high reliability

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Elias: It’s about building systems where the prompt structure itself is made robust against adversarial manipulation, which makes it much harder for attackers to exploit those perceptual gaps

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Nadia: So we’ve covered a lot of ground on this paper, and I think it’s crucial that we keep looking at how these prompt injection attacks are evolving because they are becoming increasingly subtle

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: It was fascinating seeing how the framework performs across different models like GPT-five and Llama four as it confirms that structural defenses offer broad applicability

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Conclusion: Nadia: So, we've wrapped up our discussion on "Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection," which really showed us how easy it is for attackers to bypass current AI defenses through subtle prompt injection

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Elias: Indeed, Nadia, and from a cryptographic viewpoint, seeing how they exploit perceptual asymmetry means we need to think about the assumptions in any verification process—the parameters that break down are often those things that are invisible to the human eye

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: I found what really struck me in their data was how effectively InjectDefuser reduced those attack success rates across models, specifically showing GPT-five’s vulnerability being almost entirely eliminated with that framework

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Nadia: That zero point three percent success rate is staggering; it means for all the sophisticated techniques they mapped out, InjectDefuser stops them cold, and we need to figure out how cheaply we can deploy such a robust system

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Elias: The defense framework itself is interesting because it’s not just one thing; it’s a combination of prompt hardening via UUIDs, structured URL representation, and output validation that works together to create a multi-layered barrier

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: I think the Allowlist RAG component is particularly important because it addresses those tricky failure cases where strong meta-instructions might misclassify legitimate traffic as malicious

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Nadia: So, if we can build this kind of defense, it means phishing detection systems can become much more trustworthy even against attacks that are designed to look completely normal to a user

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Elias: It certainly opens up new avenues for how we approach adversarial robustness in any complex system, not just language models

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: I think the biggest impact is that it forces us to treat the entire context of a webpage—metadata, scripts, visible content—as a unified object that needs rigorous scrutiny rather than looking at pieces in isolation

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Nadia: Exactly. So we've seen how to map out these attack surfaces and build defenses that can handle them, which is exciting for security researchers who want to understand the exploit chain

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Elias: It’s a solid paper that gives us concrete mechanisms to test against, which is something I appreciate when we're trying to understand the underlying mathematics of these vulnerabilities

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Priya: Ultimately, this research moves us closer to having AI detection that isn't just trained on known examples but is structurally resilient against novel prompt injection methods

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

Nadia: Agreed. The next challenge will be figuring out how to scale these defense mechanisms efficiently for every single model we use in production environments

Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection: .

More episodes

← Home