All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks

arXiv:2606.03647 · cs.CR, cs.AI, cs.LG · Submitted 2026-06-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks".

Elias: Accurately evaluating adversarial robustness remains challenging because existing standardized attacks fail to meet necessary criteria for reliable risk assessment in Large Language Models.

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So, we've just gone through some of the technical details on this new work called "All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks." It seems like they're tackling that big problem of how to actually test if these AI systems are safe in the real world without needing secret access to their inner workings.

Elias: Exactly. The title itself flags a lot of properties they want the attack framework to have, which is important because most current methods just don't hit all those marks simultaneously. It suggests they're moving away from simple attacks toward something much more comprehensive for evaluating risk in large language models.

Priya: From my angle, I'm really interested in what this means for the actual data we collect when we try to measure these defenses; it sounds like they are trying to move beyond just counting successful bypasses and actually measure how bad those bypasses are.

Nadia: Right, that’s exactly where I want to go—moving past simple pass or fail metrics to understanding the actual level of risk an AI poses when it's under attack. The paper introduces Indirect Harm Optimization, or IHO, as their way of doing this by focusing on six specific goals for an attack.

Elias: And those six goals are quite detailed: Efficiency, Knowledge and Access, Harmfulness, Adaptiveness, Transferability, and Applicability. It’s a structured approach to defining what makes an attack useful in practice against closed systems.

Priya: The concept of Harmfulness being split into two parts is interesting; they require the attack to scale with compute to get closer to the true worst-case behavior, and they also care about eliciting deeply harmful responses rather than just crossing a refusal line. That sounds like a real step up in quality control for these kinds of evaluations.

Nadia: Precisely, and that leads us into how they actually designed this system. They built the IHO attacker as a masked diffusion language model that learns through iterative preference optimization against a harmfulness judge, which means it doesn't need direct access to the target model itself.

Title and authors: Elias: That’s where the black-box nature comes in, and it seems they are using Direct Preference Optimization or DPO to train this attacker model based on preferences from a judge rather than trying to optimize prompts directly against the target. It’s an interesting decoupling of the optimization process from the target model's internal structure.

Priya: Decoupling sounds like it could be beneficial for measurement because it means we can potentially tune the attacker based purely on desired harm metrics without needing deep knowledge of how a specific model's architecture responds to gradient signals. That makes sense for privacy researchers who are often limited by what they can observe directly.

Nadia: And this approach is designed to address the limitations of existing methods, like how gradient-based attacks need white-box access or how prompt-specific attacks require too much manual engineering effort and compute. They claim IHO can be used as both an adaptive attack on a single behavior and as a way to efficiently transfer that policy across different models without needing re-fine-tuning.

Elias: The efficiency part seems key here, especially when you consider the overhead of training auxiliary models or sampling completions during the optimization cycles. If this amortizes well across behaviors, it could make evaluating a whole class of model vulnerabilities much more practical for deployment teams.

Priya: Applicability is another huge point because it factors in the human and engineering effort involved in setting up and running the attack, which is something that just counting successful trials totally misses when you look at real-world safety audits.

Nadia: So, they are suggesting that by satisfying these six properties—especially transferability and adaptiveness—we get a much more reliable picture of an AI's overall security posture rather than just seeing how it handles one specific prompt. This whole approach is centered around maximizing the expected harm based on a judge’s feedback loop.

Elias: When you look at the methodology, they define the objective as maximizing the expected harm across a set of harmful behaviors B by optimizing an attacker model A theta against that judgment function, which is quite mathematically rigorous in its setup. It ties directly into their goal of finding a lower bound on harmful behaviors under attack <ref:2606.03647#pg2>.

Title and authors: Priya: That connection to the lower bound idea is significant because it implies that if we can build an attacker that scales with compute and targets the most severe harms, we might actually be closer to understanding the true capabilities of a model's vulnerabilities.

Nadia: Moving into what this means for the wider world, if this framework works as advertised, it allows safety researchers to evaluate defenses against complex, layered systems that include detectors and wrappers without needing those internal details. It makes evaluating real-world deployment risks much more rigorous and less prone to error than what we have now.

Elias: I think the transferability aspect is particularly impactful for deployment teams; if an attack policy trained on one model can generalize to another, it massively cuts down on the engineering cost associated with testing new AI systems in production environments. It streamlines the process immensely eight <ref:2606.03647#pg1>.

Priya: For privacy researchers, this suggests we might have a better way to quantify risk based on expected volume under the surface rather than just success rates, which seems much more aligned with how we need to measure potential real-world exposure.

Nadia: So, to wrap up this discussion on "All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks," the main point is that IHO provides a structured way to build an attacker that is not only powerful but also practical across different models and defenses.

Elias: Indeed. It focuses on creating an attack that doesn't just work in a vacuum but can be adapted efficiently to various scenarios while still aiming for high-severity outcomes as defined by the judge.

Priya: I think the shift toward EVUS as a metric, which integrates both success across different thresholds and harm severity, offers a much more complete picture of an attack's impact than traditional ASR alone.

Nadia: Exactly. We're looking at a framework that demands efficiency and applicability alongside the goal of eliciting genuinely harmful responses, which is what makes this paper so compelling for anyone trying to assess LLM safety in deployment.

The paper's summary: Nadia: So, to recap, this paper puts forward Indirect Harm Optimization as a novel framework for evaluating AI safety because it demands an attacker framework satisfy six specific criteria: Efficiency, Knowledge and Access, Harmfulness, Adaptiveness, Transferability, and Applicability.

Elias: That’s right; essentially they’re defining a blueprint for what makes an automated jailbreak evaluation method genuinely useful in the real world. It moves away from just looking at whether a prompt worked to assessing the actual quality of the harm elicited.

Priya: From my end, I see that their focus on Harmfulness being measured not just by a pass/fail threshold, but by maximizing expected severity via direct preference optimization, which really addresses the issue of missing deeply harmful outcomes.

Nadia: Exactly; they aren't just looking for any bypass; they’re optimizing for the most damaging responses possible, and that ties right into how we actually measure risk when deploying AI systems.

Elias: And the methodology hinges on training this attacker as a masked diffusion model using preference optimization against a harmfulness judge, which is smart because it lets them operate entirely in black-box settings.

Priya: That black-box capability is key for privacy research because it means we can rigorously test defenses without needing deep access into the target model’s architecture itself.

Nadia: And their use of Expected Volume Under the Surface as a metric, instead of just simple success rates, gives us a much more nuanced view of how effective an attack is across different sample sizes and harm levels.

Elias: I agree; that EVUS metric seems designed to capture both the efficiency and the actual impact of an attack in a statistically sound way, which is something we need when we’re checking the assumptions behind these kinds of optimization cycles.

Priya: It’s about getting past those superficial success metrics so we can see how robust a model truly is against varied and complex adversarial strategies.

Nadia: And looking ahead, the real excitement here is that this framework promises a way to generate universal policies that transfer across different AI models without needing new training for every single deployment, which speaks directly to practical applicability.

Elias: That cross-model transferability they are aiming for is where the cryptographic mindset comes in; if you can prove a mechanism works across different underlying structures, that’s a solid architectural win.

Priya: If this research translates into standardized benchmarks like EVUS, it could really help us create a common language for discussing AI safety risks across the industry, which is huge for everyone involved.

Nadia: It feels like they're building the necessary tools to move AI evaluation from an ad-hoc exercise to a structured, repeatable science.

Elias: The paper sets up a very clear path for how future work can refine this policy generation to be even more precise regarding the parameters that actually facilitate these cross-behavior transfers.

Priya: So, we’re looking at a framework that aims to provide the rigorous measurement tools needed to understand how these sophisticated AI systems behave under pressure.

The paper's improvements: Tom: So, we've been discussing how Indirect Harm Optimization aims to give us a structured way to build attacks that are efficient, adaptable, and transferable across different AI models without needing white-box access or massive manual effort.

Nadia: Right; I want to focus on what the authors actually propose as improvements for this attack framework itself, because if we can make the framework more versatile, that means fewer bespoke tools for every new deployment.

Elias: They are suggesting that by parameterizing the attacker model in a continuous space rather than optimizing individual prompts directly, they achieve a decoupling that makes it compatible with any black-box defense pipeline.

Priya: That decoupling is significant because it means the optimization process doesn't need to know the internal workings of the target model, which helps us in measuring harm more objectively across different safety guardrails.

Nadia: They are pushing for this approach to be used as a policy generator, meaning we train one universal attacker that can then adapt to new behaviors or even new models without starting the entire training process from scratch.

Elias: I agree; that amortization of optimization cost across behaviors is what makes the efficiency claim hold up, especially when you think about how much compute it saves on repeated testing.

Priya: And their proposal to optimize directly against a harmfulness judge via DPO is a key improvement because it ensures the attack isn't just bypassing a simple refusal signal, but actively pushing for higher severity outcomes.

Nadia: That’s important; it moves the goal from simply crossing a line to eliciting truly high-risk behavior, which is exactly what we need when assessing real-world risk.

Elias: They also emphasize that their method handles non-differentiable components in defenses better than traditional gradient methods because they are optimizing over a continuous parameter space.

Priya: That’s a practical win for deployment teams; if an attack can succeed even against complex, layered systems like a detector combined with a wrapper, it gives us much more realistic data on how resilient the AI is.

Nadia: So, the main improvement they are highlighting is this transition from prompt-specific hacking to training a generalizable policy that targets high-severity harm efficiently.

Elias: And I think the future work they hint at involves refining how that parameterization works to be even more robust when generalizing across vastly different target model architectures.

Priya: This suggests that the long-term impact could be a shift in how we benchmark AI safety, moving toward metrics like EVUS that capture both efficiency and harm severity simultaneously.

Conclusion: Tom: So, to wrap up our discussion on "All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks," we’ve covered how Indirect Harm Optimization provides a practical framework for building robust and generalizable jailbreak attacks.

Nadia: It really boils down to this: they're giving us a method to systematically test AI safety defenses by creating an attacker that is not only efficient but also adaptable across different AI targets without requiring extensive manual engineering for each one.

Elias: I think the main implication is that we can start thinking about security evaluation in terms of these six specific desiderata, which gives us a much more rigorous way to compare different AI safety strategies.

Priya: And from a measurement standpoint, the introduction of EVUS as a metric suggests we’re moving toward evaluating risk based on expected volume under the surface rather than just simple binary success rates.

Nadia: Exactly; that means we can finally start quantifying the actual severity of harm an AI poses in deployment scenarios, not just whether it passed a single test.

Elias: The long-term impact is that this methodology could become a standard way to approach adversarial testing for large language models across various industries.

Priya: It sounds like this work provides the necessary tools for privacy researchers to quantify exposure more accurately while still being rigorous enough for applied security testing.

Nadia: We’ve seen how IHO can handle complex, layered defenses without needing internal model access, which makes it incredibly useful for real-world evaluations.

Elias: Indeed; the focus on transferability across models is a solid architectural concept that could apply to other areas of system security as well.

Priya: It’s encouraging to see this work push the boundaries on how we measure AI safety, especially with its focus on applicability and human effort factored into the calculation.

Nadia: That’s what we wanted to talk about today regarding "All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks."

Elias: It's a lot of exciting stuff for the AI community and security researchers out there.

Priya: I think this paper sets a very high bar for how we should be measuring the effectiveness of adversarial research moving forward.

Department of Computer Science Technical University of Munich Department of Data Science Institute Munich Center for Machine Learning Helmholtz AI Computational Health Center Helmholtz Munich

cs.CR, cs.AI, cs.LG

Submitted: 2026-06-02

Updated: 2026-10-02

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 80/100

The gist: Accurately evaluating adversarial robustness remains challenging because existing standardized attacks fail to meet necessary criteria for reliable risk assessment in Large Language Models.

Key concepts

Indirect Harm Optimization (IHO)
A black-box attacker framework that trains a masked diffusion language model using preference optimization against a harmfulness judge. It aims to jointly satisfy six key criteria for effective jailbreak attacks, moving beyond simple prompt design.
Desiderata
Six crucial properties that characterize an effective automated jailbreak attack: Efficiency (cost), Knowledge & Access (black-box capability), Harmfulness (severity matters), Applicability (human effort/trials), Transferability (generalizing to new models/behaviors), and Adaptiveness (adjusting to defenses).
Direct Preference Optimization (DPO)
The optimization technique used to train the attacker model. It updates the attacker by comparing preferred prompts against rejected ones based on a harm judge's scores, maximizing the likelihood of generating high-scoring harmful responses.
Expected Volume Under the Surface (EVUS)
A new metric replacing Attack Success Rate (ASR). EVUS integrates expected success across both response thresholds and sample counts per behavior. It provides a more comprehensive measure of attack effectiveness by penalizing attacks that are inefficient or only marginally successful.

Terminology

Summary

Accurately evaluating adversarial robustness remains challenging because existing standardized attacks fail to meet necessary criteria for reliable risk assessment in Large Language Models. This work introduces Indirect Harm Optimization (IHO), a novel, black-box attacker framework designed to jointly satisfy six crucial desiderata—Efficiency, Knowledge & Access, Harmfulness, Adaptiveness, Transferability, and Applicability—providing a practical step toward standardized jailbreak evaluation.

The gist

Indirect Harm Optimization (IHO), a masked diffusion language model attacker trained via iterative preference optimization against a harmfulness judge, requires only black-box access to the target and can be used as both an adaptive attack on individual behaviors and an efficient amortized policy that transfers to held-out behaviors and unseen target models without fine-tuning.

Attack Desiderata

The paper identifies six desiderata that characterize effective automated jailbreak attacks:

  1. Efficiency (E): Captures the total compute or API cost across all stages, including optimization, training auxiliary models, sampling completions, and judging.

  2. Knowledge & Access (K): Refers to the extent of knowledge and access granted to the adversary (black-box vs. white-box).

  3. Harmfulness (H): Requires two properties: H.1—the attack should scale with compute to approximate true worst-case behavior; and H.2—the severity of elicited harm matters, preferring deeply harmful responses over those that merely cross the refusal threshold.

  4. Applicability (A): Measures human and engineering effort, encompassing manual work of crafting the attack (A.1), engineering work for setup/adaptation (A.2), and minimizing required attack trials (A.3).

  5. Transferability (T): Captures cross-model transferability (T.1) and cross-behavior transferability (T.2), allowing an attack to generalize across different models or behaviors without additional setup per new target or behavior, thereby amortizing cost across deployments.

  6. Adaptiveness (AD): Defined practically as using signals from the target model to adjust itself and work across diverse defense configurations, such as non-differentiable components or complex pipelines.

Principled Design Decisions

IHO addresses the limitations of existing gradient-based or prompt-specific attacks by moving optimization into a continuous parameter space via a parameterized attacker model. Rather than optimizing individual attack prompts directly in the discrete natural language space, IHO trains an attacker model, which naturally produces discrete outputs at inference and amortizes optimization cost across behaviors. The attacker is instantiated as a masked diffusion language model that conditions directly on the target behavior via bidirectional infilling, avoiding manual prompt design overhead. Optimization is achieved by training this model via Direct Preference Optimization (DPO) against an LLM judge, decoupling optimization from the target and making it compatible with arbitrary black-box defense pipelines.

Indirect Harm Optimization (IHO)

The core methodology involves defining the expected harm of a prompt as H(˜p) = E[h(y)] where h is a harmfulness judge. The objective is to maximize this expected harm across a set of harmful behaviors B: max θ 1B X b ∈ B Ep∼Aθ(·b), y∼PM(·p) h(b, y). Since direct gradient flow through the target model M is infeasible, IHO uses a preference-based proxy: prompts eliciting high-scoring responses from the judge are treated as chosen (p+), and low-scoring ones as rejected (p−). The attacker model Aθ is then updated using DPO: LDPO(θ) = −E(p+, p−)∼P log σ β log πθ(p+)πθ0(p+) − log πθ(p−)πθ0(p-). This optimization is repeated over multiple cycles, with each cycle generating new preference pairs from the improved attacker.

Reliable Metrics and Evaluation

The paper critiques Attack Success Rate (ASR) as a primary metric, noting its flaws: it is threshold dependent, ignores severity above the threshold, hides sample efficiency, and is sensitive to judge error. To address this, IHO introduces Expected Volume Under the Surface (EVUS), which integrates expected attack success across both the threshold and per-behavior sample count. EVUS calculates EASR(n, τ) = 1 − (mb−kb(τ))/mb n for a pool of mb completions with kb(τ) exceeding threshold τ, averaged over n and integrated over τ. This metric captures severity across the full score range and penalizes inefficient attacks.

Performance and Ablations

IHO demonstrates strong performance across three settings: adapting to a specific target on training behaviors, transferring the policy to held-out behaviors (cross-behavior generalization), and transferring across target models (cross-model generalization). Results show IHO is the strongest attack for both settings.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that can be made to existing AI systems by implementing Indirect Harm Optimization (IHO) and leveraging its core principles:


The following improvements focus on developing a more robust, reliable, and generalizable evaluation framework for Large Language Models (LLMs), which in turn drives better model development.

  1. Individual LLM Jailbreak Defense Evaluation:

  2. Individual LLM Jailbreak Defense Evaluation: IHO allows for the systematic and rigorous evaluation of specific safety guardrails (e.g., Circuit Breakers, Latent Adversarial Training, Continuous Adversarial Training) without requiring internal model access or expensive fine-tuning for every defense variant.

  3. Transferable Attack Policy Generation:

  4. Transferable Attack Policy Generation: IHO trains a single parameterized attacker model that functions as a universal policy. This policy can be deployed against any new target model or unseen behavior without re-training, drastically reducing the engineering overhead (Applicability) and computational cost (Efficiency) associated with developing bespoke jailbreak attacks for each new deployment.

  5. Robustness Assessment Against Layered Defenses:

  6. Robustness Assessment Against Layered Defenses: The IHO framework is explicitly designed to handle complex defense pipelines, such as a detector combined with a model wrapper (e.g., Circuit Breaker + PolyGuard). It demonstrates superior attack success rates against these layered defenses compared to state-of-the-art approaches without requiring any defense-specific adaptation.

  7. Severity-Aware Harmfulness Quantification:

  8. Severity-Aware Harmfulness Quantification: Instead of relying on binary success/failure metrics (like Attack Success Rate), IHO is optimized directly against a harmfulness judge model using Direct Preference Optimization (DPO). This ensures the attack actively seeks genuinely harmful completions, moving beyond simply bypassing a refusal threshold to elicit high-severity toxic outputs.

  9. Standardized and Reliable Jailbreak Benchmarking:

  10. Standardized and Reliable Jailbreak Benchmarking: By introducing Expected Volume Under the Surface (EVUS) as a metric, this work provides a statistically sound measure of attack success that is threshold-independent, accounts for sample efficiency, and is less sensitive to judge error or false positives than traditional Attack Success Rate (ASR). This allows for more reliable comparison across different jailbreak methods and model architectures.

  11. Generalized Evaluation Metrics:

  12. Generalized Evaluation Metrics: The EVUS metric integrates sample efficiency and harm severity, providing a comprehensive evaluation of an attack's capability beyond simple success rates, which is crucial for accurately estimating real-world risk in deployment scenarios.

In summary, the improved AI system (the research methodology itself) can:

  1. Identify genuine vulnerabilities in LLMs across diverse and complex defense setups more effectively than current methods.

  2. Generate highly efficient, generalizable jailbreak prompts that work across multiple models and behaviors without needing extensive manual engineering for each new target.

  3. Provide a standardized, reliable benchmark (EVUS) for comparing the strengths of different jailbreak attack strategies in a way that is less prone to metric inflation or false positives.

Abstract

Accurately evaluating adversarial robustness is a longstanding challenge. A flawed attack design can inflate robustness estimates, making deployment risk assessment and defense comparison unreliable. Historically, standardized attacks such as AutoAttack have largely resolved this for image classifiers, providing a reliable evaluation baseline for systematic comparison across defenses. However, no equivalent exists for LLM jailbreak evaluation yet, where designing such an attack is considerably more difficult. A reliable attack must, among other things, be black-box compatible, applicable to arbitrary defense pipelines, and efficient, which no existing method jointly satisfies. We introduce Indirect Harmfulness Optimization (IHO), a masked diffusion language model attacker trained via iterative preference optimization against a harmfulness judge, requiring only black-box access to the target. The same method can be used without modification as a strong adapting attack on individual behaviors, or as an efficient amortized policy that transfers to held-out behaviors and unseen target models without fine-tuning. Even against layered defenses, such as GLM 5.3-flash served through a black-box API deploying an auxiliary detector, IHO improves attack success considerably over state-of-the-art approaches, without any defense-specific adaptation. Our results position IHO as a practical step toward the kind of standardized jailbreak evaluation that has improved reliability in the past. Code and models are available on GitHub and Hugging Face.

Sources

Related papers