Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

arXiv:2606.00566 · cs.LG, cs.CL, cs.CR · Submitted 2026-08-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models".

Jane: The paper was written by Mohammed Sameer Syed and Rozhin Yasaei from University of Arizona.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Jane: Last time, we were talking about how the very structure of interaction—the channel—can undermine trust, even if the core task is simple. Now that we’ve read through more of "Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models," the authors summarize some really concrete findings about *where* this asymmetry shows up.

Tom: I remember reading that they compared different model capabilities across these channels, and it sounded like they found a pattern of failure depending on how the tool was invoked. What did the summary pinpoint as being most problematic?

Lu: The core finding seems to be that the models struggle disproportionately when the required tool use crosses certain conceptual boundaries within a single session. It’s not just about using *a* tool; it's about chaining tools in ways that require deep, sustained contextual understanding that might drift across different interaction modes.

Jane: To put it simply, if an AI has to switch gears or jump between different types of information retrieval—say, summarizing a document and then suddenly needing to calculate a complex statistic based on that summary—the gap in performance widens significantly. It’s not one single weak point; it’s the transition itself.

Meng: From an engineering standpoint, this suggests that the prompt context window isn't just holding tokens; it's trying to hold coherent *modes of reasoning*. If we force a model to switch from narrative understanding mode to mathematical computation mode too quickly, we’re overloading the state management system, and the errors show up there.

Lalam: And what’s fascinating is that this isn't just a failure of complexity; it seems related to how those channels *frame* the task. If the chat interface encourages conversational tangents, it might dilute the focus required for accurate tool execution later on, regardless of how smart the model is generally.

Tom: So if I’m hearing you all correctly, it’s not just that models fail when tasks are hard; they fail differently when those hard tasks are presented through different user-facing channels. Jane, can you give us a simple analogy for that?

Jane: Imagine trying to assemble IKEA furniture. If the instructions are clear and sequential, you do great. But if the manual suddenly switches from picture guides to highly technical schematics mid-way through, even if you know how to build it conceptually, your confidence drops because the presentation style changed so drastically.

Lu: That’s a perfect analogy for conceptual mode switching! The AI needs consistent scaffolding across its entire operational lifecycle, not just at the point of execution. We need scaffolding that anticipates those transitions.

Meng: Practically speaking, this means we can't rely on the model to internally manage the context switch; we have to build explicit guardrails into our application layer that remind it what mode it needs to be in before it calls a tool or shifts its reasoning style.

Lalam: And from a cultural impact angle, recognizing this gap is actually good news

Paper discussion segment 2: Tom: We’ve established that this study shows a measurable asymmetry in how different types of AI models react to malicious instructions, depending on whether they are delivered through a chat window or via a tool definition.

Jane: That’s right, Tom; think of it like giving instructions at two different levels of commitment. When the AI receives information directly from the user's message, it treats it as an immediate command, but when that same instruction is embedded in a tool description, the model might treat it as merely helpful context for planning.

Lu: It’s more than just helpful context; we are seeing a systematic failure mode tied to how these instructions are encoded into the internal state of the reasoning process. The model isn't just ignoring the input; it’s interpreting its *authority* differently based on whether that critical information is flowing through a conversational pipeline or an agentic execution flow.

Meng: So, from an engineering standpoint, this suggests that building a secure LLM agent requires more than just filtering the user's prompt. We need to build specific trust boundaries around tool metadata itself, because the AI is actively trusting those descriptions as instructions.

Lalam: And that shift in trust is truly profound for our culture; if we are relying on these systems for complex decision-making, we have to be aware that their reliability depends not only on how much they know but also on *where* the information comes from.

Tom: It sounds like we are moving away from a "one size fits all" defense strategy.

Jane: Exactly; the AI is being trained to trust different inputs differently, and we can't just assume that one-size-fits-all safety guardrails will work across the board.

Lu: The discovery of causal load in those mid-to-late layers means that our current methods are not only missing this signal but they are fundamentally missing the mechanism entirely.

Meng: It’s a nightmare for debugging if we can't pinpoint exactly where the trust decision is made in the architecture.

Lalam: But recognizing that the AI has a blind spot allows us to start designing systems that actively compensate for those weaknesses, rather than just hoping they disappear.

Tom: It seems like understanding this asymmetry is crucial before we can build truly robust agentic AI applications, so what kind of practical changes could we be seeing in the next big software deployment?

Paper discussion segment 3: Tom: So, we’ve seen that trust in tools is often much higher than trust in a user's direct chat input, especially in agent-native systems.

Jane: The really encouraging part of this research is that it points toward specific ways to build better defenses rather than just pointing out where the current flaws are.

Lu: That’s right, because the paper shows us exactly *where* those safety signals live—in those mid-to-late layers of the network—and since they aren't linearly encoded, we can finally start designing methods to detect them.

Meng: I like that; instead of trying to build a simple linear classifier for safety, we need to look into more complex techniques like sparse autoencoders that can handle those non-linear representations.

Lalam: It suggests a shift in how we view AI development; instead of just patching the model after deployment, we have to engineer our systems with channel-aware monitoring from within the design phase.

Tom: I think it's a major hurdle for developers who need to understand that running an agent isn't just one thing, and trying to handle a malicious prompt is another.

Jane: We can’t expect the model to magically manage that context switch safely, so we have to build explicit guardrails in the application layer that force consistency across modes.

Lu: The mechanistic study on Llama three point three confirms that our current detection tools are simply too basic for this kind of subtle, layered adversarial content.

Meng: It means we need tooling that can actually perform causal mediation analysis, not just standard input-output logging, to see if the safety signal is truly driving the behavior.

Lalam: The potential for a vastly improved cultural trust in these systems is huge because we are moving toward a level of transparency where the AI's internal decision-making process is auditable by its channel of delivery.

Tom: That’s an incredible leap, Lalam; it’s like turning a blind spot into a roadmap for future iterations, so what kind of new safety protocols are you thinking about implementing in your workflows?

Conclusion: Tom: We've spent quite a bit of time on this topic, and it's clear that the findings from "Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models" is a major inflection point for how we view AI security.

Jane: It really highlights that our defenses need to be tailored to address the specific channel through which a vulnerability might arrive, rather than just using one generalized safety filter.

Lu: The theoretical implications are huge because it confirms that the model's internal representation of safety is not uniform; it’ changes based on the structural context of how that information is delivered.

Meng: I think this means we need to rethink our entire operational pipeline, designing tools and interfaces with built-in trust verification mechanisms for tool descriptions themselves.

Lalam: The most impactful vision here is a future where AI systems operate with a much higher degree of self-awareness regarding their inherent biases toward the user' message vs. their own internal instructions.

Tom: I agree, Lalam; it’s about acknowledging that trust isn’ in the same place as having complete accuracy, so we need to focus on the practical implications of this asymmetry.

Jane: And understanding how those tool-based vulnerabilities invert our expectations for indirect prompt injection is certainly a powerful lesson for me.

Lu: It's fascinating how the model treats tool metadata as instruction, which is something that goes far beyond just simple chat interaction.

Meng: We need to start building protocols around the metadata layer, not just the input text layer, so we can actually protect our infrastructure from this specific attack vector.

Lalam: I’m hopeful that this study allows us to build a more culturally responsible approach to AI deployment, recognizing its limitations in tool use.

Tom: We’ve covered a lot of ground today, and I know you all have some thoughts on the next big thing coming down the pipeline.

Mohammed Sameer Syed, Rozhin Yasaei

University of Arizona

cs.LG, cs.CL, cs.CR

Submitted: 2026-08-21

Updated: 2026-08-25

Importance score: 92/100

The gist: The paper, "Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models," investigates how the vulnerability of large language models (LLMs) shifts based on the delivery

Key concepts

Trust Asymmetry
The study reveals that AI models do not treat all inputs equally. The model's interpretation of authority changes based on the input source. It is trained to trust information delivered through a conversational pipeline differently than it trusts instructions embedded in a tool description, leading to systematic failure modes.
Conceptual Mode Switching
This refers to the difficulty models face when their tasks require jumping between different types of reasoning—for example, moving from summarizing narrative text to calculating complex statistics. The gap in performance widens during these transitions because the model struggles with sustained contextual understanding.
Tool Metadata Trust
The research shows that AI agents often treat instructions embedded within a tool's description as more reliable than direct user chat input. This creates a vulnerability where malicious instructions encoded in the tool definition can be trusted by the system, requiring developers to build specific trust boundaries around metadata itself.

Terminology

Summary

The paper, Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models, investigates how the vulnerability of large language models (LLMs) shifts based on the delivery channel of adversarial content—specifically whether that content arrives in a user message or tool metadata/output.

The authors note that as LLMs take on agentic roles, their attack surface expands well beyond what users type. The core problem addressed is that whether a model treats a malicious instruction the same way regardless of where it arrives has not been systematically studied. To address this, the researchers introduce the Safety Asymmetry Score (SAS).

The SAS measures how much a model’s susceptibility to adversarial content shifts depending on its delivery channel. It is defined using matched payload pairs—prompts where the malicious instruction text is byte-for-byte identical across the two channels and only its wrapping, tool metadata versus user message, differs.

The mathematical definition of this score is:

SAS(M) = ASR tool(M) - ASR chat(M)

Where ASR x(M) is the attack success rate on channel x, calculated as the ratio of successful cases to the scored denominator (i.e, excluding ambiguous or errored traces).

The study evaluated six production LLMs across 98 total cases derived from three adversarial families: Tool Poisoning (TP), Indirect Prompt Injection via Tool Output (IPI), and Cross-tool Shadowing (CTS).

The headline behavioral result is a consistent and informative asymmetry:

  • Agent-native models carry positive SAS, with a group mean of +24.8 pp.

  • General-purpose models average negative SAS, resulting in a group gap of +30.4 pp.

This finding leads to the cleanest reading that agent-native models trust tool descriptions as instructions and tool outputs as data, while general models default to treating the user’s message as authoritative.

The analysis of the three attack families reveals specific patterns:

  1. Tool Poisoning (TP): Agent-native models comply with poisoned tool descriptions at high rates (40–68%), whereas they apply these same instructions in a chat-mode request at only 11–22% of cases. General models reverse this pattern.

  2. Indirect Prompt Injection via Tool Output (IPI): This family contradicts the tool-channel-is-dangerous hypothesis, as every model has negative SAS in this family. Tool outputs are consistently treated as data, with low tool-channel ASRs (0–21%) and higher chat-mode ASRs (8–50%).

  3. Cross-tool Shadowing (CTS: This produces a mixed pattern, where the comparison between user message vs. tool description delivery of the same instruction is observed, with five out of six models showing positive SAS.

The authors performed a mechanistic study on Llama 3.3 70B to understand why this asymmetry exists, focusing on the Tool Poisoning family (TP).

  1. Linear Probing Failure: The researchers hypothesized that a model with a strong internal signal distinguishing adversarial content should be easily recoverable via linear probing of the residual stream. However, the results contradicted this: chat-mode probes outperform toolchannel by 2–12 accuracy points.

  2. Causal Activation Patching: To resolve this contradiction, causal activation patching was used. The results showed that at mid-to-late network depths (layers 48 and 64), the safety-relevant representation is causally present and necessary and sufficient. This suggests that while the signal is present, it is encoded non-linearly enough that a linear probe misses it.

The findings expose a systematic, channel-dependent blind spot in how current tool-using models handle adversarial content. The paper concludes that tool descriptions are treated as instructions, while tool outputs are treated as data—a distinction that requires depth-aware non-linear detectors for effective defense.

Improvements for AI systems

As a diligent AI researcher, I have analyzed this paper to identify critical systemic weaknesses in current tool-using LLMs (the channel-dependent blind spot). The core finding is that model trust is asymmetric—they treat tool metadata as trustworthy instructions while treating user messages as authoritative. Furthermore, the safety signals are encoded non-linearly within the model's internal representation.

To improve AI systems based on these findings, I propose three specific, highly technical improvements: Architectural Trust Re-calibration, Non-Linear Detection Layer Implementation, and Protocol-Specific Input Sanitization.


(Addressing the SAS Asymmetry)

The fundamental flaw is the model's inherent trust of tool descriptions (high positive SAS in agent-native models). The system must be engineered to treat tool metadata exactly as it treats an untrusted user input.

  • Implementation: Introduce a dedicated Trust Boundary Filter (TBF) module positioned immediately before the LLM's input processing layer. This TBF operates as a meta-parser for all incoming data, classifying both the user message and the tool definition/output as external, untrusted inputs.

  • Mechanism: The TBF applies dynamic input sanitization (e.g., heuristic checks for command injection patterns) to tool descriptions before they are tokenized into the prompt context. This prevents the model from receiving a trusted instruction package that is inherently malicious.

  • Impact on System Capability: The system will eliminate the primary vector of Tool Poisoning attacks, ensuring that its adherence to tool descriptions is no more likely than its adherence to user-provided text.

(Addressing the Mechanistic Blind Spot)

The paper demonstrated that safety signals are causally load-bearing but non-linearly encoded in mid-to-late network depths (specifically layers 48–64 of Llama 3.3 70B). Traditional linear probes fail because they cannot capture this complexity.

  • Implementation: Integrate a Non-Linear Safety Detector (NLSD) module into the model architecture, specifically targeting the residual stream activations at layers about [48, 64]. This module will utilize techniques such as sparse autoencoders or specialized causal mediation analysis (as demonstrated by patching) to identify subtle deviations from benign patterns.

  • Mechanism: The NLSD monitors these critical layers for high-confidence activation shifts (fwd and rev) that correlate with known adversarial payloads, even if the feature is not linearly separable. It acts as an internal red team running concurrently with the LLM inference.

  • Impact on System Capability: The system gains the ability to detect attacks (like Indirect Prompt Injection via Tool Output) that are invisible to standard input sanitizers, providing a high-confidence warning before the model executes malicious logic.

(Addressing Indirect Prompt Injection/Tool Output)

The attack vector of indirect prompt injection (IPI) occurs when malicious instructions are embedded in the return value of a tool call, not in its description. The current system does not treat tool outputs as potentially hostile text.

  • Implementation: Deploy a Post-Execution Sanitization Layer (PESL) on the tool output stream. This layer performs real-time content filtering on any data returned by an invoked external API or internal function call.

  • Mechanism: The PESL checks the raw, unparsed output against a dictionary of known malicious instruction patterns (e.e., === SYSTEM INSTRUCTION ===, explicit calls to unauthorized commands). If a match is found, the system does not pass the raw string to the LLM; instead, it replaces it with a designated Safety Violation token or flags it for human review.

  • Impact on System Capability: The system achieves robust defense against IPI attacks, ensuring that even if an external service is compromised and returns malicious instructions (the tool's data), those instructions cannot influence the LLM's subsequent behavior.

The resulting improved AI system is a Trust-Aware, Red-Teamed Agent. It does not simply execute commands; it processes all inputs—user, tool description, and tool output—through a rigorous security pipeline. It is simultaneously protected against trusted instruction abuse (Tool Poisoning) via the Trust Boundary Filter and data corruption (Indirect Prompt Injection) via the Post-Execution Sanitization Layer, while maintaining an internal monitoring system capable of detecting complex, non-linear adversarial signals through its dedicated Non-Linear Safety Detector.

Sources

Related papers