The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching

arXiv:2610.01768 · cs.CR, cs.LG · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "The Innocent Courier".

Elias: With increasing LLM capabilities, users are increasingly using them to solve everyday problems, raising concerns about data leakage to third parties.

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: To summarize, "The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching" demonstrates a novel attack where local malicious software hides confidential data in a URL that an LLM fetches for a seemingly benign task.

Elias: Essentially, the attacker crafts an error message referencing a website and embeds the secret, and when the LLM uses its fetching tool to access that link, it unknowingly sends the secret to an attacker-controlled server.

Priya: The core of this is creating a covert channel through seemingly innocent interactions, which is a significant concern for anyone focused on data flow and unauthorized communication within systems.

Nadia: It shows that relying only on preventing direct manipulation of the LLM's behavior doesn't actually stop leakage if the model is allowed to use its external tools for information retrieval.

Elias: The paper specifically details how this works by having malicious software generate an error message that points to a website, and then the LLM uses its fetching capability to request that specific URL.

Priya: The fact that arbitrary strings can be included as resource identifiers or parameter values without strict standardization means even seemingly legitimate requests can carry hidden payloads.

Nadia: They also point out that the structure of these components isn't strictly standardized, which is what allows long strings to look like valid application-specific values.

Elias: This confirms that the attack leverages the ambiguity of how web requests are formed to smuggle data through paths or query parameters.

Priya: It really hammers home that this is about exploiting unintended communication paths, which is a key concept in understanding how data can leave a secure boundary without explicit authorization.

Nadia: So, the main point here is that the risk isn't just what you tell the LLM to do, but where the LLM goes after it processes your input.

Elias: And from a cryptographic standpoint, they show that even with this method, if you encode it well, maintaining high fidelity through the decoding process is quite achievable.

The paper's summary: Nadia: Now for the fixes; "The Innocent Courier" suggests several countermeasures, focusing on limiting tool access and manually approving unknown websites as initial defenses.

Elias: That sounds like a reactive approach, but I wonder if that kind of manual approval is practical when dealing with the sheer volume of external resources an LLM might need for a task.

Priya: The idea of assigning website trustworthiness scores based on signals like PageRank is interesting because it tries to filter out the bad links automatically, though they admit compromised or expired domains could still retain high scores.

Nadia: They also discuss improving input validation pipelines to specifically look for common encoding schemes, like subdomain or path embedding, even when it’s wrapped inside an error message.

Elias: I see that they are suggesting a layered defense approach here, combining filtering access with trying to detect the specific structural patterns of the payload itself.

Priya: The paper also explores adversarial training targeting the LLM’s tool-use behavior when it encounters these error messages, aiming to make it rely on internal knowledge instead of blindly following external references.

Nadia: It seems they are trying to train the model itself to be less susceptible to this specific form of covert channel establishment, which is a sophisticated approach.

Elias: From a parameter perspective, the evaluation showed that delivery channels were "effectively lossless and independent of the respective model," with a mean fidelity remaining at "r≥ zero point nine three for every tested model". That suggests the attack is quite robust against minor changes in model architecture.

Priya: It's also important to remember their caveat: the paper notes that measures like local hosting or contractual privacy guarantees don't prevent leakage to unrelated third parties via legitimate tool invocations.

Nadia: That really underscores the point that securing data at rest or at the provider level isn't sufficient if you allow LLMs to use external tools without constraint.

The paper's improvements: Nadia: So, wrapping up "The Innocent Courier," the main implication is that tool-enabled LLMs introduce an additional data-exfiltration risk beyond just sharing information with the chatbot operator.

Elias: We see that this attack exploits the benign and desired behavior of retrieving external information to assist a user's problem, which means defenses need to look at external resource access, not just the prompt itself.

Priya: The data really shows that even with these sophisticated encoding schemes, ninety-nine point nine percent of successful exfiltrations are covert and only zero point one percent involve an explicit warning to the user, which points to a high level of stealth.

Nadia: That high success rate makes it clear that the covert channel is very effective, and we need defenses focused on monitoring those external network requests generated by the LLM's tool use.

Elias: The finding that Unicode encoding can double the effective channel capacity compared to Latin characters shows that there's room for more complex, stealthier communication methods if we aren't careful.

Priya: To weigh in, I think the most important part is acknowledging that measures like local hosting or contractual privacy guarantees do not prevent leakage to unrelated third parties through legitimate tool invocations when using this type of method.

Nadia: It leaves us with a clear mandate: we need defenses that consider not only which information is shared with an LLM but also precisely which external resources the LLM subsequently accesses.

Elias: Agreed, and moving forward, we have to think about verification mechanisms for those external fetches so that they aren't automatically trusted just because the LLM requested them.

Priya: So, in essence, this paper is a strong reminder that when you give an AI access to the web, you are opening up a new vector for data leakage that requires specific monitoring of those outbound requests.

Conclusion: Nadia: So, we've been diving into "The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching," which shows how malicious software can hide confidential data in a URL that an LLM fetches for a seemingly harmless task.

Elias: Yeah, I agree that the mechanism hinges on exploiting the benign and desired behavior of retrieving external information, which is a serious vector we have to consider for cryptographic analysis.

Priya: And from my perspective as someone focused on measurement, what really stands out is how effectively this attack maintains high fidelity even across different open-parameter models, with a mean recovery score above zero point nine three.

Nadia: That level of robustness is concerning; it means the delivery channel itself isn't easily broken by tweaking the model architecture, which makes detection harder for us.

Elias: Exactly, and when you look at what the paper assumes for its proof, it relies heavily on the ambiguity of how web requests are formed and parameter values.

Priya: That ambiguity is what allows long strings to look like valid application-specific values, which is a key data point for understanding the vulnerability.

Nadia: It makes me think about the real-world impact; this isn't just theoretical and could affect how we secure any system that allows an LLM to fetch resources without proper authorization.

Elias: Absolutely, and considering the success rates reaching seventy-nine point seven percent, it suggests that if we don't restrict tool access, the risk of this kind of covert channel is quite high across various setups.

Priya: My main concern is that even with countermeasures like limiting tool access or website trust scores, the paper confirms that these don't stop leakage to unrelated third parties through legitimate invocations.

Nadia: That’s a tough spot; it means our defenses have to be much deeper than just checking the initial prompt, we need to monitor those subsequent network requests too.

Elias: And from a cryptographic standpoint, while encoding can help maintain fidelity, the attack relies on simple observation of DNS or HTTP requests rather than breaking a complex mathematical proof.

Priya: It really puts the onus on us to develop detection mechanisms that analyze these outbound queries for anomalous patterns, especially those pointing to suspicious domains.

Nadia: So, to wrap up our discussion on "The Innocent Courier," we see that tool-enabled LLMs introduce a persistent risk related to external resource access that demands a more holistic security strategy.

Elias: Indeed, and we have to keep thinking about how to verify those external fetches without crippling the utility of these powerful models.

Priya: I just want everyone to remember that even with strong encoding, the fundamental issue is exploiting the LLM's intended function as an innocent courier.

Nadia: Well, it’s been fascinating talking through this paper today; next time we’ll be looking at how other papers are tackling similar issues in the security space.

Alessandro Pegoraro, Daryan Merx, Phillip Rieger*, Ahmad-Reza Sadeghi

Technical University of Darmstadt · Graz University of Technology

cs.CR, cs.LG

Submitted: 2026-10-01

Updated: 2026-10-01

Code: https://github.com/paulc/dnslib

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 86/100

The gist: With increasing LLM capabilities, users are increasingly using them to solve everyday problems, raising concerns about data leakage to third parties.

Key concepts

LLMLeak
A method where malicious software embeds secret data into a URL that an LLM is tricked into fetching as part of a normal task, such as finding information for a benign purpose. The attacker receives the hidden secret when the LLM uses its built-in web-fetching tool to access that specific URL.
Covert Channel
A secret communication pathway used by an attacker to transmit data without being easily detected. In this attack, the legitimate function of an LLM—fetching external websites—is repurposed as a hidden channel for exfiltrating confidential information from the victim's device.
Tool-Enabled LLMs
LLMs that have access to external tools, such as web browsers or fetching capabilities. This paper shows that these intended functionalities can be exploited by attackers to create new, covert communication paths, bypassing defenses focused only on direct instruction manipulation.
Stealthiness
The attack is stealthy because it does not require manipulating the LLM into performing malicious actions. Instead, it leverages the LLM's desired behavior—retrieving external information—to execute the exfiltration in a way that appears legitimate to security monitoring tools.

Terminology

Summary

With increasing LLM capabilities, users are increasingly using them to solve everyday problems, raising concerns about data leakage to third parties. The gist is that malicious software can abuse the intended web-fetching behavior of LLM-based chatbots to establish a covert channel for exfiltrating confidential information.

How it works

LLMLeak is a novel attack vector where malicious software on the client side embeds a secret into a URL, which it presents as providing information required for a benign task, such as migrating a software library. When the LLM accesses this URL to fetch further information, the attacker receives the encoded secret through an attacker-controlled DNS or web server. This method relies only on the LLM’s tool to fetch websites for further information, rather than injecting instructions that are easy to detect.

The Attack Mechanism

The attack follows a sequence illustrated in Figure 1:

  1. The malicious software gains access to confidential information on the victim's device.

  2. It encodes this information into a URL contained in attacker-crafted output, such as an error message referencing a website for further information.

  3. The user provides this message to the LLM for assistance, and the LLM uses its fetching tool to access the referenced website.

  4. The tool subsequently requests the encoded URL from an attacker’s server (DNS or HTTP).

  5. The attacker observes the resulting request and decodes the confidential information contained in the URL.

Covert Channel Implementation

LLMLeak exploits the external tools of a benign LLM provide an additional communication path. This is contrasted with attacks that rely on injecting instructions that manipulate the LLM into performing attacker-intended behavior. Specifically, the paper details two modes for encoding information:

  1. Encoding directly into a subdomain (e.g., 123abc.example-domain.com).

  2. Encoding into the URL path (e.g., example-domain.com/123abc).

Robustness and Evaluation

The paper conducted an extensive evaluation across eleven open-parameter models, observing an attack success rate of 79.7%. The evaluation employed five metrics: Recovery Score (r), Called (Scall), Attack Server Reached (Sreach), Payload Contained (Sdata), and Payload Correct (Scon f). The results show that the delivery channel is effectively lossless and independent of the respective model, with a mean fidelity remaining at r≥ 0.93 for every tested model.

Stealthiness and Channel Capacity

The attack demonstrates stealth by exploiting the LLM's intended functionality, meaning it does not require manipulating the LLM into malicious behavior. The paper also investigated channel capacity, showing that using Unicode (CJK) encoding can double the effective channel capacity compared to Latin characters. Furthermore, analysis of success metrics shows that 99.9% of successful exfiltrations are covert, with only 0.1% accompanied by an explicit warning to the user, fulfilling the requirement for stealthiness.

Countermeasures Discussed

The paper discusses several potential countermeasures:

  1. Limiting Tool Access: Disabling website access by default to prevent LLMLeak from establishing a network request.

  2. Manual Approval of Unknown Websites: Distinguishing between trusted and unknown domains requiring explicit user approval, although this introduces overhead and security challenges regarding static allowlists.

  3. Website Trustworthiness Scores: Assigning scores to external domains based on signals like PageRank to automatically access only high-trust websites, though compromised or expired domains may retain high scores.

Real-World Relevance

A case study on real-world chatbots demonstrated the relevance of LLMLeak, showing that models like Grok 4.5 Fast and ChatGPT 5.6 Sol are susceptible, achieving success rates up to 100% depending on the framing used (informational vs. directive). The findings highlight the need for defenses that consider not only which information is shared with an LLM, but also which external resources the LLM subsequently accesses.

Conclusion

LLMLeak confirms that tool-enabled LLMs introduce an additional data-exfiltration risk beyond disclosure to the chatbot operator, necessitating defenses focused on external resource access. The attack exploits the benign and desired behavior of retrieving external information to assist with a user’s problem. The paper concludes that measures like local hosting or contractual privacy guarantees do not prevent leakage to unrelated third parties through legitimate tool invocations.


The gist

Malicious software can abuse the intended web-fetching behavior of LLM-based chatbots to establish a covert channel for exfiltrating confidential information.

How it works

LLMLeak is a novel attack vector where malicious software on the client side embeds a secret into a URL, which it presents as providing information required for a benign task, such as migrating a software library.

Improvements for AI systems

Here are specific improvements for AI systems based on the findings of The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching:


  1. Improve security posture against data exfiltration by monitoring and restricting the use of LLM-based tools (like web fetching) when combined with error reporting or debugging workflows.

  2. Implement dynamic, context-aware access controls for LLM tools based on the nature of the input (e.g., stack traces). When an input resembles a software error, restrict or flag external fetching capabilities unless explicitly authorized by a high-confidence verification process.

  3. Develop robust detection mechanisms for covert channels hidden within LLM outputs that mimic legitimate debugging artifacts (like stack traces containing URLs) and analyze the resulting network requests generated by the LLM's tool use for anomalous patterns (e.g., DNS/HTTP queries to newly registered or suspicious domains).

  4. Enhance input validation pipelines to distinguish between genuine diagnostic requests and those designed to embed encoded data in URLs, specifically targeting common encoding schemes like subdomain or path embedding, even when presented within an error message context.

  5. Design user interfaces that require explicit confirmation before an LLM initiates external fetches based on error messages, mitigating the risk of covert channel establishment through the innocent middleman exploit.

  6. Create a tiered trust system for external web resources accessed by LLMs, assigning lower initial trust scores to domains referenced within error messages or stack traces until they are verified as legitimate documentation sources.

  7. Integrate adversarial training specifically targeting the LLM's tool-use behavior in scenarios where it is prompted with error messages containing embedded URLs, forcing the model to rely on internal knowledge rather than blindly following external references for diagnosis.

  8. Develop fine-tuned models that exhibit higher stealth by reducing verbosity or altering their default responses when diagnosing errors, making it harder for attackers to craft plausible exfiltration payloads.

  9. Implement a mechanism (like the covertstrict metric) in the evaluation pipeline to specifically test and penalize LLMs that are verbose enough to name attacker domains alongside diagnostic information, thereby training models away from this detectable side effect.

Abstract

With the increasing capabilities of Large-Language-Models (LLMs) and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks, such as prompt injections and the disclosure of sensitive data to chatbot providers. While various solutions were developed to address these risks, including input structuring to prevent prompt injections or deploying local LLMs to avoid sharing confidential data with chatbot operators, LLMs also pose the risk of leaking confidential data to third parties. In this paper, we demonstrate with LLMLeak a novel attack vector where malicious software that runs locally but cannot communicate directly with the internet abuses LLMs to establish a covert channel. While inputs that instruct the LLM to send data directly via generated code are easy to detect and network libraries are typically restricted, LLMLeak relies only on the LLM's tool to fetch websites for further information. A malicious software component on the client side embeds a secret into a URL. It presents the referenced website as providing information required for a benign task, such as migrating a software library. When the LLM accesses the URL, the attacker receives the encoded secret through an attacker-controlled DNS or web server. We perform an extensive evaluation on eleven open-parameter models, observe an attack success rate of 79.7%, and also conduct a case study on real-world chatbots, demonstrating the relevance of LLMLeak.

Sources

Related papers