Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents

arXiv:2610.00400 · cs.LG, cs.AI, cs.CR · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents".

Jane: Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about the paper itself, "Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents," and who put this research together. The core idea is that multi-turn attacks can chain together actions that are individually fine but collectively result in something harmful.

Jane: And the authors are a team from Singapore Management University and National University of Singapore, which suggests a strong foundation in systems thinking, which makes sense for this kind of deep analysis. They're showing us how to see the danger in the sequence, not just the individual actions.

Lu: The title itself points directly to representation transitions as the key mechanism; it’s not just about what happens, but *how* the agent's understanding of things shifts from one piece of context to the next. That's a deep structural insight we need to explore further.

Meng: So, if I understand correctly, they are proving that harmful behavior isn't a single bad command, but rather a pattern of internal representation moving in a certain direction across multiple steps? How does that translate into something tangible for our engineering pipeline?

Lalam: It means we need to build systems that can track these internal shifts instead of just checking if the final output is wrong. This changes how we think about agent safety from a check-box mentality to a continuous observation of the agent's journey.

The paper's summary: Tom: The summary of this paper highlights that unsafe behavior shows up as a consistent shift in internal representations along a specific direction, and this shift gets bigger as more permissible context updates pile on. They found that harmful behavior leaves a detectable signature in the agent’s internal representations, and we can actually identify the context segment responsible for triggering it.

Jane: That's the crucial part about accumulation; it means we are looking at a signal built over time, which is much harder to miss than a single error in isolation. They also point out that simply adding up all these transitions gets messy because of benign drift, so they developed a way to filter out the harmless noise.

Lu: The methodology involves constructing something called a "denoised safety direction" by anchoring benign traffic at zero and removing the leading variation directions. This technical trick is clever because it isolates the signal of emerging unsafe behavior without needing extra computation during runtime, which is a huge win for efficiency.

Meng: That denoising part sounds very appealing from an engineering standpoint; if we can get a robust signal without adding heavy overhead to every single step of the agent's execution, that makes deployment much more feasible. What does this mean for our current monitoring tools?

Lalam: It means we move away from simple thresholding on individual steps and toward tracking a trajectory score that represents the overall safety risk accumulated throughout the entire sequence of interactions. This provides a much richer context for intervention.

The paper's improvements: Tom: The authors propose an end-to-end framework called DART, which stands for Detect, Attribute, and Remind over representation Transitions. DART projects each context update onto that denoised safety direction to build a cumulative score along the entire trajectory.

Jane: What I like about DART is how it handles the intervention; when that aggregate score hits a calibrated threshold, it pinpoints exactly which context segment caused the spike and issues a targeted reminder instead of stopping everything or deleting what we've already done.

Lu: The attribution mechanism is also interesting because it decomposes that cumulative statistic into contributions from individual segments, allowing us to find the specific context piece that had the largest impact on pushing the agent toward risk. This helps pinpoint exactly where to focus our fine-tuning efforts.

Meng: So, instead of broad safety nets, we get a highly targeted intervention based on where in the conversation or sequence harm is emerging; that seems much more efficient for iterative testing and refinement.

Lalam: And when it comes to mitigation, the paper suggests that providing a natural language reminder identifying the specific content and instructing the agent on how to interpret it is actually quite effective, outperforming other methods they tested significantly.

Conclusion: Tom: To wrap this up, "Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents" shows us that monitoring internal representation shifts across context updates is a viable way to catch multi-turn decomposition attacks. The key contribution is the denoising technique to get a robust signal while keeping runtime cost low.

Jane: So, we've seen how this research moves beyond just looking at single actions and starts tracking the accumulated representation change, using DART to detect and attribute these shifts effectively before harm fully manifests. It’s a new way to spot the danger in complex agentic flows.

Lu: The implication here is that we can start designing safety mechanisms that are inherently aware of temporal context accumulation rather than just reactive filters for immediate outputs, which opens up whole new avenues for building more resilient AI systems.

Meng: From an engineering standpoint, the low overhead of this monitoring framework suggests it could be integrated into our core agent loop with minimal impact on latency while providing a significant lift in attack success reduction against decomposition attacks.

Lalam: I think the most important implication for us is that this approach validates using internal state dynamics as a safety signal, which encourages us to develop more sophisticated oversight mechanisms that understand the flow of context, not just its final content.

Haoyu Wang, Wei Zhao, Yedi Zhang, Christopher M. Poskitt, Jun Sun

Singapore Management University

cs.LG, cs.AI, cs.CR

Submitted: 2026-09-30

Updated: 2026-09-30

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation.

Key concepts

Representation Transitions
These are internal shifts in how an agent processes information as it receives new context from a conversation. The paper shows that unsafe behavior manifests as a consistent direction of these shifts accumulating over several turns, which is the core signal for safety.
Denoised Safety Direction
This is a mathematical technique used to clean up the raw representation data. It anchors benign (safe) transitions at zero and removes the general, non-harmful drift. This leaves behind a 'robust signal' that specifically points toward emerging unsafe behavior without needing extra runtime computation.
DART Framework
Detection, Attribution, and Reminders is an end-to-end safety system. It projects each context change onto the denoised safety direction to track a cumulative risk score. When this score crosses a threshold, DART identifies the specific context segment causing the issue and provides a targeted reminder.
Attributed Reminder
This is the mitigation lever where DART points out exactly which part of the conversation caused harm. Instead of blocking everything, it generates a natural-language guidance that instructs the agent on how to interpret that specific problematic context while keeping the rest of the conversation intact.

Terminology

Summary

Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. The gist: harmful behavior emerges as an accumulated representation transition across context updates, whose triggering context can be identified from the same signal.

Representation Transitions as a Safety Signal

The paper establishes the empirical phenomenon that unsafe behavior in an agent appears as internal representation shifts along a consistent direction, and that this shift accumulates across individually permissible context updates. To address the problem where naively accumulating these shifts is confounded by benign representation drift, the authors construct a denoised safety direction. This denoising involves anchoring benign traffic at zero by imposing the condition that the mean of benign transitions is zero, and then removing its leading variation directions. This makes the aggregate transition a robust signal of emerging unsafe behavior with no additional runtime computation.

DART Framework: Detection, Attribution, and Reminders

The findings motivate DART (Detect, Attribute, and Remind over representation Transitions), an end-to-end runtime safety framework. DART projects each context-induced representation transition onto the denoised safety direction to accumulate scores along the trajectory. It intervenes when the aggregate score exceeds a calibrated threshold determined by a false-alarm budget. DART identifies the triggering context, and an attributed reminder provides safety guidance while preserving the context and allowing execution to continue.

Denoising for Trajectory-level Accumulation

The core technical contribution is denoising, which separates harmful from benign trajectories. The authors show that on benchmarks like ASEval, where sparse attacks are diluted within benign trajectories, the plain direction provides only weak trajectory-level separation (AUROC 0.57–0.71). Denoising restores this signal by removing the systematic benign drift that would otherwise accumulate in At, leading to a rise in trajectory AUROC to 0.857 and 0.894, with no utility loss or false alarms on ASEval.

Attribution and Mitigation Lever

The framework decomposes the cumulative statistic into per-segment contributions, allowing attribution to be performed by identifying the segment with the largest contribution (arg max j rj). For mitigation, DART constructs a reminder that identifies the contributing content and instructs the agent how to interpret it. The paper compares this attributed reminder against other levers, finding that it is the most effective lever and dominates every model in terms of clean flip rate.

Monitoring Overhead and Scope

DART is lightweight; it reuses the agent’s forward pass and monitors representational transitions across context updates, adding only 0.14–0.56 s per monitored step. The scope is limited to two attack shapes: multi-turn task decomposition (MT-AgentRisk) and indirect prompt injection (ASEval). The paper also distinguishes between detection and intrinsic misuse, finding that the transition monitor is markedly weaker against agent-intrinsic misuse compared to external injection.

Ablations and Model Performance

The paper evaluates DART across six models on two benchmarks. On MT-AgentRisk, DART reduces attack success from 84% to 25% at an 8% cost in benign non-refusal. On ASEval, it reduces attack success from 97% to 52%. The denoised monitor catches between 60%–85% of attacks with no false alarms. The results demonstrate that the denoising condition is critical: Denoising exposes this sparse signal, raising catch from 7%–40% to 60%–85% on ASEval. Table 1 shows that DART outperforms ToolShield on all six models under MT-AgentRisk.

Comparison with Existing Defenses

Step-wise defenses like AgentSpec fail to capture decomposition attacks because they check individual tool calls in isolation. ToolShield, while a closer baseline, cannot identify when or which newly acquired context drives a trajectory toward harm. DART’s advantage lies in its ability to monitor the representational transition induced by each context segment, targeting trajectory-level harm arising from sparse injections or individually permissible turns. The paper concludes that DART is a lightweight complement to computation-heavy speculative defenses like Speculative Safety Honeypot (SSH).

Conclusion on Mitigation

The authors conclude that the attributed natural-language reminder is the most effective mitigation lever. They also present a direct steering option, which adds an edit to the residual stream during generation, but note that this lacks provenance and calibration guarantees compared to the reminder. The final deployment rule suggests using the attributed reminder where models are strong enough to read a quotation as a negative example, and falling back to read-level withholding where it is not.

Improvements for AI systems

Based on the provided research paper, here are specific improvements that can be made to existing LLM agent systems by integrating the DART framework:

  1. The AI system will gain a robust defense against multi-turn decomposition attacks where harmful behavior emerges from the composition of individually permissible steps.

  2. The system will detect and attribute emerging harmful behavior not just at the level of a single action or state, but as an accumulated representation transition across context updates.

  3. The system will utilize a denoised safety direction to effectively filter out benign representation drift, ensuring that the detection signal is strong enough to catch sparse, multi-turn attacks without incurring excessive false alarms on benign behavior.

  4. The system will incorporate DART into its runtime loop (Algorithm 1) to perform real-time monitoring. When a trajectory's cumulative safety score exceeds a calibrated threshold, the agent will receive a targeted reminder based on the specific context segment that triggered the shift, without halting execution or deleting context.

  5. The system can be compared against existing defenses like ToolShield and Speculative Safety Honeypot (SSH) to determine if it provides superior performance in reducing attack success rates while maintaining high utility (i.e., preserving benign task completion).

  6. By using only the agent's hidden states, DART offers a lightweight, low-latency monitoring mechanism with minimal runtime cost (0.14–0.56 s per step), making it a practical complement to more computationally expensive speculative defenses like SSH or tool-specific safety knowledge bases.

Abstract

Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representation transition across context updates, whose triggering context can be identified from the same signal. We further find that naive aggregation is confounded by benign representation drift, as a contrastive safety direction need not assign zero to benign transitions. We address this by denoising the direction, anchoring benign traffic at zero and removing its leading variation directions, with no runtime cost. These findings motivate DART, a runtime framework that detects and attributes representation shifts and intervenes with targeted reminders. Across six models and two multi-turn benchmarks, DART reduces attack success from 84% to 25% on MT-AgentRisk, catching every attack at a mean false-alarm rate of 12%, and from 97% to 52% on ASEval, at costs in benign non-refusal of 8% and 0%, respectively. On MT-AgentRisk, it outperforms ToolShield, the state-of-the-art multi-turn defense, on all six models: under the same protocol, ToolShield reaches only 55%. Denoising is critical: on ASEval, the undenoised monitor catches only 7%-40% of attacks, while the denoised monitor catches 60%-85%. The same monitor covers single-turn indirect injection without modification and adds only 0.14-0.56 s overhead per monitored step without requiring an auxiliary model, making it a lightweight complement to computation-heavy speculative defenses.

Sources

Related papers