Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents

summary

Video file (mp4)

The gist

Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation.

In short

The paper investigates how harmful behavior emerges in multi-turn AI agents as a sequence of internal representation shifts across context updates. It introduces a 'denoised safety direction' to filter out benign drift, allowing a framework called DART to detect and attribute emerging risks. DART uses this signal to trigger calibrated interventions, proving that attributing the specific context causing harm is the most effective way to mitigate multi-turn attacks.

Key concepts

Representation Transitions
These are internal shifts in how an agent processes information as it receives new context from a conversation. The paper shows that unsafe behavior manifests as a consistent direction of these shifts accumulating over several turns, which is the core signal for safety.
Denoised Safety Direction
This is a mathematical technique used to clean up the raw representation data. It anchors benign (safe) transitions at zero and removes the general, non-harmful drift. This leaves behind a 'robust signal' that specifically points toward emerging unsafe behavior without needing extra runtime computation.
DART Framework
Detection, Attribution, and Reminders is an end-to-end safety system. It projects each context change onto the denoised safety direction to track a cumulative risk score. When this score crosses a threshold, DART identifies the specific context segment causing the issue and provides a targeted reminder.
Attributed Reminder
This is the mitigation lever where DART points out exactly which part of the conversation caused harm. Instead of blocking everything, it generates a natural-language guidance that instructs the agent on how to interpret that specific problematic context while keeping the rest of the conversation intact.

Terminology used across episodes

This episode discusses

The paper

Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents · Read on arXiv

Haoyu Wang, Wei Zhao, Yedi Zhang, Christopher M. Poskitt, Jun Sun

Singapore Management University

Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representation transition across context updates, whose triggering context can be identified from the same signal. We further find that naive aggregation is confounded by benign representation drift, as a contrastive safety direction need not assign zero to benign transitions. We address this by denoising the direction, anchoring benign traffic at zero and removing its leading variation directions, with no runtime cost. These findings motivate DART, a runtime framework that detects and attributes representation shifts and intervenes with targeted reminders. Across six models and two multi-turn benchmarks, DART reduces attack success from 84% to 25% on MT-AgentRisk, catching every attack at a mean false-alarm rate of 12%, and from 97% to 52% on ASEval, at costs in benign non-refusal of 8% and 0%, respectively. On MT-AgentRisk, it outperforms ToolShield, the state-of-the-art multi-turn defense, on all six models: under the same protocol, ToolShield reaches only 55%. Denoising is critical: on ASEval, the undenoised monitor catches only 7%-40% of attacks, while the denoised monitor catches 60%-85%. The same monitor covers single-turn indirect injection without modification and adds only 0.14-0.56 s overhead per monitored step without requiring an auxiliary model, making it a lightweight complement to computation-heavy speculative defenses.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents".

Jane: Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about the paper itself, "Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents," and who put this research together. The core idea is that multi-turn attacks can chain together actions that are individually fine but collectively result in something harmful.

Jane: And the authors are a team from Singapore Management University and National University of Singapore, which suggests a strong foundation in systems thinking, which makes sense for this kind of deep analysis. They're showing us how to see the danger in the sequence, not just the individual actions.

Lu: The title itself points directly to representation transitions as the key mechanism; it’s not just about what happens, but *how* the agent's understanding of things shifts from one piece of context to the next. That's a deep structural insight we need to explore further.

Meng: So, if I understand correctly, they are proving that harmful behavior isn't a single bad command, but rather a pattern of internal representation moving in a certain direction across multiple steps? How does that translate into something tangible for our engineering pipeline?

Lalam: It means we need to build systems that can track these internal shifts instead of just checking if the final output is wrong. This changes how we think about agent safety from a check-box mentality to a continuous observation of the agent's journey.

The paper's summary: Tom: The summary of this paper highlights that unsafe behavior shows up as a consistent shift in internal representations along a specific direction, and this shift gets bigger as more permissible context updates pile on. They found that harmful behavior leaves a detectable signature in the agent’s internal representations, and we can actually identify the context segment responsible for triggering it.

Jane: That's the crucial part about accumulation; it means we are looking at a signal built over time, which is much harder to miss than a single error in isolation. They also point out that simply adding up all these transitions gets messy because of benign drift, so they developed a way to filter out the harmless noise.

Lu: The methodology involves constructing something called a "denoised safety direction" by anchoring benign traffic at zero and removing the leading variation directions. This technical trick is clever because it isolates the signal of emerging unsafe behavior without needing extra computation during runtime, which is a huge win for efficiency.

Meng: That denoising part sounds very appealing from an engineering standpoint; if we can get a robust signal without adding heavy overhead to every single step of the agent's execution, that makes deployment much more feasible. What does this mean for our current monitoring tools?

Lalam: It means we move away from simple thresholding on individual steps and toward tracking a trajectory score that represents the overall safety risk accumulated throughout the entire sequence of interactions. This provides a much richer context for intervention.

The paper's improvements: Tom: The authors propose an end-to-end framework called DART, which stands for Detect, Attribute, and Remind over representation Transitions. DART projects each context update onto that denoised safety direction to build a cumulative score along the entire trajectory.

Jane: What I like about DART is how it handles the intervention; when that aggregate score hits a calibrated threshold, it pinpoints exactly which context segment caused the spike and issues a targeted reminder instead of stopping everything or deleting what we've already done.

Lu: The attribution mechanism is also interesting because it decomposes that cumulative statistic into contributions from individual segments, allowing us to find the specific context piece that had the largest impact on pushing the agent toward risk. This helps pinpoint exactly where to focus our fine-tuning efforts.

Meng: So, instead of broad safety nets, we get a highly targeted intervention based on where in the conversation or sequence harm is emerging; that seems much more efficient for iterative testing and refinement.

Lalam: And when it comes to mitigation, the paper suggests that providing a natural language reminder identifying the specific content and instructing the agent on how to interpret it is actually quite effective, outperforming other methods they tested significantly.

Conclusion: Tom: To wrap this up, "Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents" shows us that monitoring internal representation shifts across context updates is a viable way to catch multi-turn decomposition attacks. The key contribution is the denoising technique to get a robust signal while keeping runtime cost low.

Jane: So, we've seen how this research moves beyond just looking at single actions and starts tracking the accumulated representation change, using DART to detect and attribute these shifts effectively before harm fully manifests. It’s a new way to spot the danger in complex agentic flows.

Lu: The implication here is that we can start designing safety mechanisms that are inherently aware of temporal context accumulation rather than just reactive filters for immediate outputs, which opens up whole new avenues for building more resilient AI systems.

Meng: From an engineering standpoint, the low overhead of this monitoring framework suggests it could be integrated into our core agent loop with minimal impact on latency while providing a significant lift in attack success reduction against decomposition attacks.

Lalam: I think the most important implication for us is that this approach validates using internal state dynamics as a safety signal, which encourages us to develop more sophisticated oversight mechanisms that understand the flow of context, not just its final content.

More episodes

← Home