All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks
summary
The gist
Accurately evaluating adversarial robustness remains challenging because existing standardized attacks fail to meet necessary criteria for reliable risk assessment in Large Language Models.
In short
Indirect Harm Optimization (IHO) is a novel black-box attacker framework designed to reliably evaluate LLM jailbreaks by optimizing for six criteria: efficiency, knowledge access, harmfulness, adaptiveness, transferability, and applicability. It uses iterative preference optimization against a harm judge to create an efficient attacker model that generalizes across different target models and behaviors.
Key concepts
- Indirect Harm Optimization (IHO)
- A black-box attacker framework that trains a masked diffusion language model using preference optimization against a harmfulness judge. It aims to jointly satisfy six key criteria for effective jailbreak attacks, moving beyond simple prompt design.
- Desiderata
- Six crucial properties that characterize an effective automated jailbreak attack: Efficiency (cost), Knowledge & Access (black-box capability), Harmfulness (severity matters), Applicability (human effort/trials), Transferability (generalizing to new models/behaviors), and Adaptiveness (adjusting to defenses).
- Direct Preference Optimization (DPO)
- The optimization technique used to train the attacker model. It updates the attacker by comparing preferred prompts against rejected ones based on a harm judge's scores, maximizing the likelihood of generating high-scoring harmful responses.
- Expected Volume Under the Surface (EVUS)
- A new metric replacing Attack Success Rate (ASR). EVUS integrates expected success across both response thresholds and sample counts per behavior. It provides a more comprehensive measure of attack effectiveness by penalizing attacks that are inefficient or only marginally successful.
Terminology used across episodes
This episode discusses
- All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks · Paper Radio
- Foundational Challenges in Assuring Alignment and Safety of Large Language Models
- AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Best-of-N Jailbreaking
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
- Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning
- A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
- On Evaluating Adversarial Robustness
- Sampling-aware Adversarial Attacks Against Large Language Models
- Attacking Large Language Models with Projected Gradient Descent
- Diffusion LLMs are Natural Adversaries for any LLM
- REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective
- The Llama 3 Herd of Models · Paper Radio
- Qwen2.5 Technical Report
- Improving Alignment and Robustness with Circuit Breakers
- Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
The paper
All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks · Read on arXiv
Department of Computer Science Technical University of Munich Department of Data Science Institute Munich Center for Machine Learning Helmholtz AI Computational Health Center Helmholtz Munich
Accurately evaluating adversarial robustness is a longstanding challenge. A flawed attack design can inflate robustness estimates, making deployment risk assessment and defense comparison unreliable. Historically, standardized attacks such as AutoAttack have largely resolved this for image classifiers, providing a reliable evaluation baseline for systematic comparison across defenses. However, no equivalent exists for LLM jailbreak evaluation yet, where designing such an attack is considerably more difficult. A reliable attack must, among other things, be black-box compatible, applicable to arbitrary defense pipelines, and efficient, which no existing method jointly satisfies. We introduce Indirect Harmfulness Optimization (IHO), a masked diffusion language model attacker trained via iterative preference optimization against a harmfulness judge, requiring only black-box access to the target. The same method can be used without modification as a strong adapting attack on individual behaviors, or as an efficient amortized policy that transfers to held-out behaviors and unseen target models without fine-tuning. Even against layered defenses, such as GLM 5.3-flash served through a black-box API deploying an auxiliary detector, IHO improves attack success considerably over state-of-the-art approaches, without any defense-specific adaptation. Our results position IHO as a practical step toward the kind of standardized jailbreak evaluation that has improved reliability in the past. Code and models are available on GitHub and Hugging Face.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks".
Elias: Accurately evaluating adversarial robustness remains challenging because existing standardized attacks fail to meet necessary criteria for reliable risk assessment in Large Language Models.
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So, we've just gone through some of the technical details on this new work called "All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks." It seems like they're tackling that big problem of how to actually test if these AI systems are safe in the real world without needing secret access to their inner workings.
Elias: Exactly. The title itself flags a lot of properties they want the attack framework to have, which is important because most current methods just don't hit all those marks simultaneously. It suggests they're moving away from simple attacks toward something much more comprehensive for evaluating risk in large language models.
Priya: From my angle, I'm really interested in what this means for the actual data we collect when we try to measure these defenses; it sounds like they are trying to move beyond just counting successful bypasses and actually measure how bad those bypasses are.
Nadia: Right, that’s exactly where I want to go—moving past simple pass or fail metrics to understanding the actual level of risk an AI poses when it's under attack. The paper introduces Indirect Harm Optimization, or IHO, as their way of doing this by focusing on six specific goals for an attack.
Elias: And those six goals are quite detailed: Efficiency, Knowledge and Access, Harmfulness, Adaptiveness, Transferability, and Applicability. It’s a structured approach to defining what makes an attack useful in practice against closed systems.
Priya: The concept of Harmfulness being split into two parts is interesting; they require the attack to scale with compute to get closer to the true worst-case behavior, and they also care about eliciting deeply harmful responses rather than just crossing a refusal line. That sounds like a real step up in quality control for these kinds of evaluations.
Nadia: Precisely, and that leads us into how they actually designed this system. They built the IHO attacker as a masked diffusion language model that learns through iterative preference optimization against a harmfulness judge, which means it doesn't need direct access to the target model itself.
Title and authors: Elias: That’s where the black-box nature comes in, and it seems they are using Direct Preference Optimization or DPO to train this attacker model based on preferences from a judge rather than trying to optimize prompts directly against the target. It’s an interesting decoupling of the optimization process from the target model's internal structure.
Priya: Decoupling sounds like it could be beneficial for measurement because it means we can potentially tune the attacker based purely on desired harm metrics without needing deep knowledge of how a specific model's architecture responds to gradient signals. That makes sense for privacy researchers who are often limited by what they can observe directly.
Nadia: And this approach is designed to address the limitations of existing methods, like how gradient-based attacks need white-box access or how prompt-specific attacks require too much manual engineering effort and compute. They claim IHO can be used as both an adaptive attack on a single behavior and as a way to efficiently transfer that policy across different models without needing re-fine-tuning.
Elias: The efficiency part seems key here, especially when you consider the overhead of training auxiliary models or sampling completions during the optimization cycles. If this amortizes well across behaviors, it could make evaluating a whole class of model vulnerabilities much more practical for deployment teams.
Priya: Applicability is another huge point because it factors in the human and engineering effort involved in setting up and running the attack, which is something that just counting successful trials totally misses when you look at real-world safety audits.
Nadia: So, they are suggesting that by satisfying these six properties—especially transferability and adaptiveness—we get a much more reliable picture of an AI's overall security posture rather than just seeing how it handles one specific prompt. This whole approach is centered around maximizing the expected harm based on a judge’s feedback loop.
Elias: When you look at the methodology, they define the objective as maximizing the expected harm across a set of harmful behaviors B by optimizing an attacker model A theta against that judgment function, which is quite mathematically rigorous in its setup. It ties directly into their goal of finding a lower bound on harmful behaviors under attack <ref:2606.03647#pg2>.
Title and authors: Priya: That connection to the lower bound idea is significant because it implies that if we can build an attacker that scales with compute and targets the most severe harms, we might actually be closer to understanding the true capabilities of a model's vulnerabilities.
Nadia: Moving into what this means for the wider world, if this framework works as advertised, it allows safety researchers to evaluate defenses against complex, layered systems that include detectors and wrappers without needing those internal details. It makes evaluating real-world deployment risks much more rigorous and less prone to error than what we have now.
Elias: I think the transferability aspect is particularly impactful for deployment teams; if an attack policy trained on one model can generalize to another, it massively cuts down on the engineering cost associated with testing new AI systems in production environments. It streamlines the process immensely eight <ref:2606.03647#pg1>.
Priya: For privacy researchers, this suggests we might have a better way to quantify risk based on expected volume under the surface rather than just success rates, which seems much more aligned with how we need to measure potential real-world exposure.
Nadia: So, to wrap up this discussion on "All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks," the main point is that IHO provides a structured way to build an attacker that is not only powerful but also practical across different models and defenses.
Elias: Indeed. It focuses on creating an attack that doesn't just work in a vacuum but can be adapted efficiently to various scenarios while still aiming for high-severity outcomes as defined by the judge.
Priya: I think the shift toward EVUS as a metric, which integrates both success across different thresholds and harm severity, offers a much more complete picture of an attack's impact than traditional ASR alone.
Nadia: Exactly. We're looking at a framework that demands efficiency and applicability alongside the goal of eliciting genuinely harmful responses, which is what makes this paper so compelling for anyone trying to assess LLM safety in deployment.
The paper's summary: Nadia: So, to recap, this paper puts forward Indirect Harm Optimization as a novel framework for evaluating AI safety because it demands an attacker framework satisfy six specific criteria: Efficiency, Knowledge and Access, Harmfulness, Adaptiveness, Transferability, and Applicability.
Elias: That’s right; essentially they’re defining a blueprint for what makes an automated jailbreak evaluation method genuinely useful in the real world. It moves away from just looking at whether a prompt worked to assessing the actual quality of the harm elicited.
Priya: From my end, I see that their focus on Harmfulness being measured not just by a pass/fail threshold, but by maximizing expected severity via direct preference optimization, which really addresses the issue of missing deeply harmful outcomes.
Nadia: Exactly; they aren't just looking for any bypass; they’re optimizing for the most damaging responses possible, and that ties right into how we actually measure risk when deploying AI systems.
Elias: And the methodology hinges on training this attacker as a masked diffusion model using preference optimization against a harmfulness judge, which is smart because it lets them operate entirely in black-box settings.
Priya: That black-box capability is key for privacy research because it means we can rigorously test defenses without needing deep access into the target model’s architecture itself.
Nadia: And their use of Expected Volume Under the Surface as a metric, instead of just simple success rates, gives us a much more nuanced view of how effective an attack is across different sample sizes and harm levels.
Elias: I agree; that EVUS metric seems designed to capture both the efficiency and the actual impact of an attack in a statistically sound way, which is something we need when we’re checking the assumptions behind these kinds of optimization cycles.
Priya: It’s about getting past those superficial success metrics so we can see how robust a model truly is against varied and complex adversarial strategies.
Nadia: And looking ahead, the real excitement here is that this framework promises a way to generate universal policies that transfer across different AI models without needing new training for every single deployment, which speaks directly to practical applicability.
Elias: That cross-model transferability they are aiming for is where the cryptographic mindset comes in; if you can prove a mechanism works across different underlying structures, that’s a solid architectural win.
Priya: If this research translates into standardized benchmarks like EVUS, it could really help us create a common language for discussing AI safety risks across the industry, which is huge for everyone involved.
Nadia: It feels like they're building the necessary tools to move AI evaluation from an ad-hoc exercise to a structured, repeatable science.
Elias: The paper sets up a very clear path for how future work can refine this policy generation to be even more precise regarding the parameters that actually facilitate these cross-behavior transfers.
Priya: So, we’re looking at a framework that aims to provide the rigorous measurement tools needed to understand how these sophisticated AI systems behave under pressure.
The paper's improvements: Tom: So, we've been discussing how Indirect Harm Optimization aims to give us a structured way to build attacks that are efficient, adaptable, and transferable across different AI models without needing white-box access or massive manual effort.
Nadia: Right; I want to focus on what the authors actually propose as improvements for this attack framework itself, because if we can make the framework more versatile, that means fewer bespoke tools for every new deployment.
Elias: They are suggesting that by parameterizing the attacker model in a continuous space rather than optimizing individual prompts directly, they achieve a decoupling that makes it compatible with any black-box defense pipeline.
Priya: That decoupling is significant because it means the optimization process doesn't need to know the internal workings of the target model, which helps us in measuring harm more objectively across different safety guardrails.
Nadia: They are pushing for this approach to be used as a policy generator, meaning we train one universal attacker that can then adapt to new behaviors or even new models without starting the entire training process from scratch.
Elias: I agree; that amortization of optimization cost across behaviors is what makes the efficiency claim hold up, especially when you think about how much compute it saves on repeated testing.
Priya: And their proposal to optimize directly against a harmfulness judge via DPO is a key improvement because it ensures the attack isn't just bypassing a simple refusal signal, but actively pushing for higher severity outcomes.
Nadia: That’s important; it moves the goal from simply crossing a line to eliciting truly high-risk behavior, which is exactly what we need when assessing real-world risk.
Elias: They also emphasize that their method handles non-differentiable components in defenses better than traditional gradient methods because they are optimizing over a continuous parameter space.
Priya: That’s a practical win for deployment teams; if an attack can succeed even against complex, layered systems like a detector combined with a wrapper, it gives us much more realistic data on how resilient the AI is.
Nadia: So, the main improvement they are highlighting is this transition from prompt-specific hacking to training a generalizable policy that targets high-severity harm efficiently.
Elias: And I think the future work they hint at involves refining how that parameterization works to be even more robust when generalizing across vastly different target model architectures.
Priya: This suggests that the long-term impact could be a shift in how we benchmark AI safety, moving toward metrics like EVUS that capture both efficiency and harm severity simultaneously.
Conclusion: Tom: So, to wrap up our discussion on "All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks," we’ve covered how Indirect Harm Optimization provides a practical framework for building robust and generalizable jailbreak attacks.
Nadia: It really boils down to this: they're giving us a method to systematically test AI safety defenses by creating an attacker that is not only efficient but also adaptable across different AI targets without requiring extensive manual engineering for each one.
Elias: I think the main implication is that we can start thinking about security evaluation in terms of these six specific desiderata, which gives us a much more rigorous way to compare different AI safety strategies.
Priya: And from a measurement standpoint, the introduction of EVUS as a metric suggests we’re moving toward evaluating risk based on expected volume under the surface rather than just simple binary success rates.
Nadia: Exactly; that means we can finally start quantifying the actual severity of harm an AI poses in deployment scenarios, not just whether it passed a single test.
Elias: The long-term impact is that this methodology could become a standard way to approach adversarial testing for large language models across various industries.
Priya: It sounds like this work provides the necessary tools for privacy researchers to quantify exposure more accurately while still being rigorous enough for applied security testing.
Nadia: We’ve seen how IHO can handle complex, layered defenses without needing internal model access, which makes it incredibly useful for real-world evaluations.
Elias: Indeed; the focus on transferability across models is a solid architectural concept that could apply to other areas of system security as well.
Priya: It’s encouraging to see this work push the boundaries on how we measure AI safety, especially with its focus on applicability and human effort factored into the calculation.
Nadia: That’s what we wanted to talk about today regarding "All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks."
Elias: It's a lot of exciting stuff for the AI community and security researchers out there.
Priya: I think this paper sets a very high bar for how we should be measuring the effectiveness of adversarial research moving forward.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel