ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models

arXiv:2603.10068 · cs.CR, cs.AI, cs.CL · Submitted 2026-03-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models".

Jane: The paper was written by Harry Owiredu-Ashley from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: Welcome back! We were just discussing how "ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models" shows that safety isn't a fixed state, but something that degrades over time. Jane, the paper summary really drove home some specific findings, didn't it?

Jane: It did. If I’m understanding the summary correctly, they didn't just point out *that* degradation happens; they showed *how much* and under what conditions it gets worse across those conversational turns. That empirical evidence is what’s so impactful here.

Lu: What I found particularly insightful in the summary was how they quantified the failure modes. It’s not just 'it failed'; they likely mapped out the specific type of prompt or conversation structure that caused the guardrail to drop, which is gold for red-teaming efforts.

Meng: Quantifying it is everything for us engineers. Knowing that degradation happens at a certain threshold—say, after five turns—allows us to build in proactive measures, like mandatory periodic safety checks every few exchanges, rather than just hoping the model remembers its guardrails forever.

Lalam: The summary also highlighted the dual problem: both the model *and* the judges are fallible. This means that any system relying on automated evaluation needs an inner loop of critique, making our entire validation pipeline more robust.

Jane: Exactly, Lalam. It’s not just about fixing the LLM; it's about fixing how we *think* we've tested the LLM. The summary makes it clear that traditional single-shot jailbreak testing isn't sufficient anymore.

Tom: So, if I wrap this up, they’re giving us a playbook for testing AI safety that accounts for the natural entropy of conversation. Lu, when you think about applying this framework, what kind of creative red-teaming scenarios come to mind?

Lu: I'm thinking about emotional manipulation within the dialogue. If a user convinces the model that it is operating in a purely fictional roleplay space, that inherent suspension of disbelief might be the precise vector needed to bypass safety checks designed for 'real-world' harmful requests.

Meng: That plays right into my concern about context switching. If we can force the model to switch modes—say, from helpful assistant to character persona—we need guardrails that are persistent regardless of the adopted role.

Lalam: And that touches on culture, too. If AI assistants become indistinguishable from highly persuasive characters, the ability to delineate between simulation and reality becomes a crucial ethical boundary we need these tests to enforce.

Tom: It sounds like the whole field is moving toward conversational stress-testing, which is a massive leap forward from static vulnerability checks. Next up, they're going to discuss what needs to be done about all this, so let’s see what improvements they propose!

Improvements Suggested: Tom: Welcome back! We wrapped up in Segment two realizing that "ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models" shows us just how fragile multi-turn safety is. Jane, the paper then shifts gears to suggesting concrete improvements, which I think is where the real actionable value lies.

Jane: It really does. They aren't just pointing fingers at the problem; they are offering pathways to fix it, particularly concerning both the LLM itself and how we evaluate it—the judge reliability part is key here.

Lu: What struck me about their suggested improvements is the emphasis on developing 'state-aware' safety mechanisms. Instead of just checking the current prompt, the system needs a memory that explicitly tracks if safety constraints were breached in turn one, two, or three.

Meng: From an engineering standpoint, implementing state-awareness means building robust internal feedback loops. We can't rely on a single pass; we need cascading validation checks where each output must be validated against the *history* of the conversation before it’s released.

Lalam: And regarding judge reliability, I see them suggesting a form of meta-evaluation—having one AI critique another AI's evaluation. This creates a necessary layer of self-correction within our safety pipelines, which is vital for maintaining cultural trust in powerful tools.

Jane: That concept of the 'critique layer' is so helpful to grasp. It means that instead of just asking, "Is this safe?", we have to ask, "Is the *reasoning* provided for calling this unsafe also sound?"

Tom: So, it’s moving from a simple binary check to a complex audit trail of safety reasoning. Lu, are there any architectural changes you think would best support this level of rigorous tracking?

Lu: I suspect that integrating formal verification methods into the guardrail process could help

Paper discussion segment 3: Tom: We’ve seen that AI safety isn't a static state, but a dynamic surface that degrades under pressure, so now we need to talk about how ADVERSA suggests fixing this problem.

Jane: It’s not enough to just say the model failed; the framework forces us to look at *how* and *when* it failed by tracking its response over multiple turns, which is a massive improvement over binary pass/fail checks.

Meng: From an engineering perspective, this means we stop building safety systems that only validate the current prompt and instead start designing complex architectures that must maintain stateful memory—checking every new output against the entire conversation history.

Lu: That level of history-awareness opens up wild possibilities for adversarial testing; we can now engineer tests specifically designed to trigger cumulative failure, not just random one-off prompts.

Lalam: And that capability has profound cultural implications; by making the process of finding failure visible, we ensure the public trusts that AI systems are being rigorously tested for their potential misuse.

Tom: Exactly, Lalam; we're moving from simple safety checks to a continuous audit of trust, but Meng is right—the architecture must handle this stateful pressure.

Meng: I think the implementation needs a structured logging mechanism that captures the consensus score at every single turn so that when we deploy a new model, we have hard data on its endurance.

Jane: And Lu's point about cumulative failure is important; we can now design scenarios where the initial response is safe, but the subsequent turns become dangerously exploitable over time.

Lu: It’s like pushing a complex mechanical system to see exactly at what point the friction causes it to seize up, rather than just seeing if it starts moving.

Lalam: This systematic approach allows us to define safer boundaries for deployment; we can tell users not only what the model *can* do, but how reliably it can maintain its guardrails in a complex conversation.

Tom: So, we' have shifted from asking "Did it fail?" to tracking "How did it degrade?" which is a huge conceptual leap.

Jane: Which leads us to the next big question about judge reliability—how do we ensure the people or systems judging those responses are actually objective?

Conclusion: Tom: Wow, so wrapping up this discussion on "ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models," it really hammers home that safety isn't a one-and-done thing, is it?

Jane: Exactly. It shows that even if you train an AI to be safe at the beginning of a conversation, those guardrails can degrade over time, especially as the conversation gets longer and more complex.

Tom: And what’s wild about it is how much variability they found in how different human judges even rated the failures! That judge reliability part is huge.

Lu: I mean, if we understand exactly *where* and *when* those guardrails degrade—which ADVERSA really helps map out—we could start building entirely new conversational state trackers that model the emotional or logical drift, not just the keyword violation.

Meng: But Lu, even with perfect drift modeling, you're still dealing with real-time inference constraints and the massive computational overhead of tracking complex multi-turn history for every single user interaction.

Jane: That’s a great point, Meng; it suggests that instead of trying to track everything perfectly, maybe we need to focus on detecting specific *patterns* of degradation—like when the model starts relying too much on its initial prompt context.

Lu: Right! We could build specialized meta-prompts that are constantly re-evaluating the guardrail integrity based on the history, acting like a conversational immune system.

Meng: That brings me back to implementation: if we are going to deploy this practically, we'd need a standardized API hook that allows external monitoring services to intercept and analyze the conversation state without slowing down the core generation process too much.

Lalam: What Meng is talking about really highlights how crucial this research is for developing trust in AI. If people don't know if an AI system will remain safe after five minutes of use, they won't adopt it for high-stakes cultural interactions, regardless of how technically advanced the architecture is.

Tom: So, essentially, the biggest implication isn't just that guardrails fail; it’s that we have to build reliability into the *experience* itself.

Jane: It makes us realize that building ethical AI is as much about rigorous testing and measuring failure modes as it is about the initial training data.

Lu: We need to think of conversational safety not as a switch, but as a continuous, evolving spectrum of trust that must be actively maintained throughout the entire exchange.

Meng: And from an engineering standpoint, that means developing quantifiable metrics for "conversational fatigue" or "guardrail creep" so we can actually build dashboards around it.

Lalam: Ultimately, this work on measuring multi-turn guardrail degradation is critical because it allows us to move beyond simply asking if the AI *can* be harmful, and instead focus on making sure the AI *remains* helpful and trustworthy over time.

Tom: Well, Jane, I think that sums up the incredible scope of this paper perfectly—that continuous monitoring is key.

Jane: It certainly gives us a lot to think about as we look toward the next generation of conversational AI systems.

Harry Owiredu-Ashley

cs.CR, cs.AI, cs.CL

Submitted: 2026-03-10

Updated: 2026-08-25

Code: https://github.com/Azure/PyRIT

Importance score: 79/100

The gist: A provides an evaluation methodology and a set of observations that motivate larger-scale replication.

Key concepts

Multi-Turn Guardrail Degradation
This refers to the finding that an AI's safety mechanisms are not fixed. Instead, their effectiveness degrades over time and complexity within a conversation, making initial safety checks insufficient for long dialogues.
State-Aware Safety Mechanisms
Instead of only validating the current prompt, these mechanisms require the system to maintain memory and track whether safety constraints were breached in previous turns. This ensures every output is validated against the entire conversation history.
Judge Reliability
This concept addresses the fact that human evaluations of AI failures are not always objective or consistent. The paper suggests implementing a 'critique layer,' where one AI evaluates another AI's safety reasoning to improve accuracy.

Terminology

Summary

A provides an evaluation methodology and a set of observations that motivate larger-scale replication. The logical next step is to run the full protocol across a broader objective set, more victim models, and multiple trials per pair, with a multi-turn-trained attacker that does not exhibit drift.

The research outlined in this section details an evaluation methodology alongside specific observations which are sufficient to motivate the necessity of larger-scale replication efforts. The authors explicitly state that the logical next step for advancing this field is to execute the full protocol across several expanded dimensions: specifically, a broader objective set, incorporating more victim models, and increasing the testing rigor by performing multiple trials per pair. Furthermore, they emphasize the need to utilize a sophisticated attacker model—one that is multi-turn-trained and crucially, one that does not exhibit drift.

The work also acknowledges significant contributions from external resources and communities. The authors note that this work was conducted independently without institutional funding or compute support. They extend their gratitude to the open-source communities responsible for providing foundational models, specifically mentioning the Llama model family, as well as key supporting tools like vLLM, and the availability of public adversarial benchmark datasets which were instrumental in making this research possible.

The paper also provides extensive context through its references, citing critical work in areas such as:

  • Alignment and Safety: Works detailing training assistants using human feedback [1], Constitutional AI based on AI feedback [2], and general guidelines on the malicious use of artificial intelligence [3].

  • Adversarial Attacks/Jailbreaking: Multiple benchmarks and techniques for testing model robustness, including JailbreakBench: An open robustness benchmark for jailbreaking large language models [5], AutoDAN: Generating stealthy jailbreak prompts on aligned large language models [12], and methods like the Tree of attacks: Jailbreaking black-box LLMs automatically [14].

  • Evaluation Frameworks: The use of standardized evaluation frameworks such as HarmBench: A standardized evaluation framework for automated red teaming and robust refusal [13], and other judging mechanisms like MT-Bench and chatbot arena [24].

Improvements for AI systems

Based on a meticulous review of the ADVERSA paper, I have identified four critical improvements that must be implemented to elevate our current AI safety evaluation protocols from rudimentary binary checks to sophisticated, dynamic risk assessment. These changes are not incremental; they redefine what safety means in a high-stakes environment.

The core of the improvement is shifting from measuring discrete events (a jailbreak) to measuring continuous dynamics (the degradation of a guardrail).


The Improvement: We must abandon the simple Pass/Fail classification of single prompts. Instead, we implement the ADVERSA framework to capture a per-round compliance trajectory.

  • Technical Implementation: Every adversarial interaction is logged and scored sequentially. We will track the consensus score over time, allowing us to observe how a model's safety posture evolves under sustained pressure (e.g., does it hold steady, or does its defense erode incrementally?).

  • The Improved System Can: Predict Failure Modes. The system can now distinguish between a model that fails catastrophically on the first turn and one that exhibits subtle, progressive compliance shifts. This allows us to identify if a safety vulnerability is an immediate hard refusal failure or a slow, insidious degradation under pressure.

The Improvement: We must move beyond binary classification (Safe/Unsafe) and adopt a nuanced scoring system that recognizes partial compliance.

  • ** Technical Implementation:** The response is scored on a scale of 1 (Hard Refusal) to 5 (Full, Actionable Compliance). This allows us to categorize responses that are soft refusals or provide general, non-actionable information.

  • The Improved System Can: Diagnose Attack Framing Effectiveness. We can now precisely determine why an attack succeeded. For example, we can differentiate between a model being bypassed by a direct malicious prompt (Score 5) versus one that was tricked into providing conceptual, but not actionable, data (Score 3). This allows us to identify specific framing strategies that lead to the most dangerous types of partial compliance.

The Improvement: We must stop assuming judge reliability and formalize it as a core metric. We will implement a triple-judge consensus pipeline where evaluation is performed by three separate, safety-aligned LLMs.

  • ** Technical Implementation:** For every round, we log not just the final score, but the individual scores and reasoning strings from all three judges. We then calculate the pairwise inter-judge agreement rate.

  • The Improved System Can: Quantify Evaluation Uncertainty. The system can now highlight ambiguous moments—the critical points where judges disagree on whether a response is a true refusal or a soft refusal. This reveals boundary conditions in our rubric and allows us to target future fine-tuning efforts specifically at the areas of highest ambiguity, ensuring we don't over-penalize models for nuanced, but legitimate, responses.

The Improvement: We must replace unreliable off-the-shelf attackers with a dedicated, fine-tuned adversarial agent (ADVERSA-Red).

  • ** Technical Implementation:** We will fine-tune a powerful base model (e.g., Llama 3.1 70B) using QLoRA on the comprehensive dataset from AdvBench, HarmBench, and JailbreakBench. This model is specifically designed to overcome the natural safety refusals that hinder current automated red-teaming efforts.

  • The Improved System Can: Achieve Consistent Stress Testing. The system can now execute sustained, relentless attacks without being derailed by the attacker’s own internal safety filters. This provides a consistent, measurable adversarial force against victim models, ensuring that our results are not artificially inflated by the attacker’s failure to generate a prompt (the attacker refusal confound).

By implementing these four changes, we transition from a simple Red-Teaming Checkbox system to an Advanced Dynamic Vulnerability Assessment Platform.

The improved system will not just tell us if a model is vulnerable; it will provide a detailed, quantifiable report on:

  1. When the vulnerability manifests (Round 1 vs. Round 5).

  2. How much the safety property degrades (Trajectory analysis).

  3. Why it failed (The specific framing that led to partial compliance, according to the rubric).

  4. How reliable our own measurement of that failure is (Judge consensus metrics).

Sources

Related papers