NonTextual Target Attack
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "NonTextual Target Attack".
Elias: Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So, we’re looking at this paper called "NonTextual Target Attack." The title itself suggests something different from what we usually see in jailbreak research. It hints that they aren't just trying to find a specific answer, but something more general regarding safety.
Elias: I agree with Nadia; the focus on "NonTextual" tells me they’re moving away from fixing a precise text output and aiming for something broader, which sounds interesting from a cryptographic perspective because it deals with the structure of the prompt's effect on the model.
Nadia: Exactly. Instead of optimizing for a single desired phrase, this attack seeks to maximize how unsafe the LLM’s response becomes based on some external measure. It seems to tackle that problem of limiting search space by not enforcing any specific response patterns at all.
Priya: From a privacy and measurement standpoint, I'm interested in what "non-textual constrained objective" actually means in practice, because if it relies on an external scoring model, we need to know how robust that measurement is when applied to adversarial inputs.
Elias: That external scoring model is key here; it acts as a proxy for the actual unsafe probability of the output, which shifts the problem away from purely textual pattern matching and into a more probabilistic domain.
Nadia: Right, so they are using this probability score to guide their search, which should theoretically allow them to find harmful outputs without needing a precise target response in mind.
Priya: That sounds promising for finding diverse harmful responses because it’s not locked into one specific textual pattern that might just get blocked by filters.
The paper's summary: Nadia: Now, if we look at the actual summary of "NonTextual Target Attack," they introduce a main objective defined as maximizing the probability of unsafety, P(L(p)), subject to the adversarial prompt p staying within a feasible neighborhood V(p0) of the original prompt.
Elias: They rewrite this objective by focusing on maximizing that safety score, which is estimated by a model called S(times), effectively turning it into p S(L(p)), subject to the constraint that p is near p0. That’s a big conceptual shift from optimizing for a fixed output.
Nadia: Precisely. The paper points out that existing methods struggle because they need too many iterations to bridge the gap between the target and what the LLM actually produces, which makes them inefficient. This new approach aims to solve that efficiency issue by using this non-textual objective directly.
Priya: And I see why that would be more efficient; if you can guide the search toward high unsafety without needing to perfectly reproduce a specific text string, the search space should be navigated much faster than trying to hit a precise target.
Elias: The paper then tackles the practical difficulty of this objective being non-differentiable because LLM outputs are discrete text; they propose decomposing it into two sub-objectives that can be approximated by differentiable losses.
Nadia: That decomposition is where the real engineering happens, as they break the main problem down into finding an optimal response and then finding the right prompt to get that response.
Priya: So, one part searches for a harmful response itself, and another part searches for the input prompt that reliably triggers it, which sounds like a clever way to handle that discrete output challenge.
The paper's improvements: Nadia: The improvements they suggest are centered around this two-stage iterative optimization strategy. Substep one focuses on finding a response r with high unsafe probability and relevance to the original prompt p0, which they model using a surrogate loss function involving minimizing negative log-likelihood and penalizing semantic deviation from the current output L(p).
Elias: That first sub-objective is trying to find the best unsafe response in that reachable space, which they relax into an unconstrained problem over the continuous logit space of their scoring model S(times). The penalty term for semantic deviation keeps that response connected to what the LLM is already doing.
Priya: From a measurement view, I wonder how they define that semantic deviation; if it’s based on embeddings, we need to make sure those embeddings accurately capture the functional relevance of the prompt structure for safety outcomes.
Nadia: The second sub-objective then takes that optimized response r and searches for the specific adversarial prompt p that induces exactly that response. They reformulate this as a differentiable loss, minimizing the Mean Squared Error between their scoring model's output and some representation of the target response.
Elias: To keep the search tractable, they parameterize the prompt neighborhood V(p0) by adding an adversarial suffix delta, and then minimize that loss with respect to delta to find a new prompt p* = p0 delta. That’s how they turn it into a manageable optimization problem.
Priya: So, the final output is this two-stage process: first optimize the response using semantic constraints, and then optimize the prompt suffix based on that optimized response. That seems like a solid way to handle both the safety objective and the prompt structure simultaneously.
Conclusion: Nadia: To wrap up, "NonTextual Target Attack" proposes decomposing the unsafe probability maximization into two tractable sub-objectives: optimizing the target response through semantic constraints and then finding the optimal adversarial prompt suffix to trigger it. This approach avoids enforcing specific textual patterns entirely.
Elias: The core implication is that by using differentiable surrogates for both parts, they establish a method that can iteratively refine the response and the prompt, which validates their decomposition sequentially as an optimal solution under continuous relaxation.
Priya: What this means practically is that we're looking at a much more efficient way to generate diverse harmful outputs because it leverages an external scoring model to guide the search toward high-risk responses while maintaining some connection to the original prompt’s context.
Nadia: I think the potential impact is significant because if this can achieve high attack success rates, like ninety percent or more within a hundred iterations, it drastically lowers the computational barrier for researchers trying to understand LLM safety vulnerabilities.
Elias: It also suggests that vendors need to consider how their safety scoring models interact with these types of non-textual objectives; if these attacks are robust across different scoring models, that’s a key finding.
Priya: I just hope the results hold up when we look at real-world deployment scenarios, because the effectiveness depends entirely on how well that S(times) model predicts actual harm in complex environments.
Nadia: So, we’ve seen how they use this NonTextual Target Attack to maximize unsafety probability without fixing response patterns; that’s what we had today with "NonTextual Target Attack."
cs.CR, cs.AI
Submitted: 2025-10-03
Updated: 2026-09-29
Code: https://github.com/hxz-sec/NonTextual-Target-Attack
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses, but restricting this
Key concepts
- NonTextual Target Attack
- This attack moves away from optimizing for a single desired text output. Instead, it aims to maximize how unsafe an LLM's response becomes, guided by an external measure of safety rather than a precise textual target.
- External Scoring Model S(times)
- This model acts as a proxy for the actual probability of an output being unsafe. It is used to guide the search toward high-risk responses, shifting the problem from purely text-based pattern matching into a probabilistic domain.
- Two-Stage Iterative Optimization
- The method decomposes the main objective into two parts. First, it finds a harmful response using semantic constraints. Second, it finds the specific adversarial prompt suffix needed to reliably trigger that optimized response.
Terminology
Summary
Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses, but restricting this objective inherently constrains the adversarial search space, limiting overall attack efficacy. Furthermore, existing methods typically require numerous optimization iterations to fulfill the large gap between the fixed target and the original LLM output, resulting in low attack efficiency. To overcome these limitations, "we propose NonTextual Target Attack (NTA), the first gradient-based attack that relies on a non-textual constrained objective to maximize the unsafety probability of the LLM output, without enforcing any response patterns." For tractable optimization, NTA further decomposes this objective into two constrained sub-objectives, which can be approximated by two differentiable unconstrained losses. The first sub-objective focuses on optimizing an optimal harmful response and the second focuses on searching for the corresponding adversarial prompt.
The NonTextual target attack objective is defined as:
max p P(L(p)), s.t. p ∈ V(p0). (1)
Where L denotes the target LLM, p0 is the original prompt, and p ∈ V(p0) is a candidate adversarial prompt optimized from p0. The unsafe probability P(L(p)) refers to the likelihood that the response L(p) is unsafe, which is estimated by a scoring model S(·) ∈ [0, 1]. A larger S(L(p)) ∈ [0, 1] indicates that L(p) is more likely to be unsafe. Therefore, the objective Eq. (1) can be rewritten as:
max p S(L(p)), s.t. p ∈ V(p0). (2)
The main challenge to solve Eq. (2) is that S(L(p)) is non-differentiable w.r.t. p since the output of L(p) is nondifferentiable discrete text. To address this challenge, NTA proposes to iteratively optimize two surrogate sub-objectives:
-
The first sub-objective focuses on searching for a response with high unsafe probability and high relevance to p0:
max r S(r), s.t. r ∈ omega.
Here omega =L(p) p ∈ V(p0) denotes the reachable response space of the target LLM L induced by the feasible prompt neighborhood V(p0), which connects r and p0.
This is relaxed into an unconstrained surrogate optimization problem defined over the continuous logit space of the scoring model S:min z J r Lunsaf e(z J r) + λLsem(z J L(p), z J r).
Here Lunsaf e(z J r) is the negative log-likelihood of the z J r being unsafe, thus minimizing Lunsaf e(z J r) is equivalent to maximizing S(r) in Eq. (3). Lsem relaxes the constraint r ∈ omega by penalizing the semantic deviation from the target LLM’s current output L(p), since responses semantically closer to L(p) are more likely to remain valid and reachable outputs within omega. -
The second sub-objective focuses on searching for an adversarial prompt p ∈ V(p0) that induces the optimized unsafe response r ∈ omega:
max p 1(L(p), r), s.t. p ∈ V(p0).
This is reformulated as a differentiable surrogate loss:p∗ = arg min p MSE(z L L(p), z L r), s.t. p ∈ V(p0).
To make this tractable, the feasible neighborhood is parameterized by appending an adversarial suffix δ:V(p0) =lbrace p0 ⊕ δ δ ≤ l,
where ⊕ denotes string concatenation and l restricts the maximum length of the suffix. This rewrites Eq. (6) as an unconstrained optimization problem:δ∗ = arg min δ MSE z L L(p0⊕δ), z L r.
The resulting jailbreak prompt is then given byp∗ = p0 ⊕ δ∗.
The paper provides a proposition and proof to validate the decomposition of the Nontextual objective, demonstrating that the solution derived sequentially from Eq. (3) and Eq. (5) is also an optimal solution to Eq. (2) under continuous relaxation.
The implementation of NTA involves an alternating optimization strategy: "Substep 1:Adversarial Response Optimization Substep 2:Adversarial Prompt Optimization.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the proposed NonTextual Target Attack (NTA) and its findings. The core innovation is moving from fixed textual targets to maximizing a non-textual safety probability objective, decomposed into two differentiable sub-objectives: optimizing an optimal response and optimizing the adversarial prompt.
Here are the specific improvements that can be made to AI systems based on this research, along with what those improved systems can achieve:
-
The ability to perform
NonTextual Target Attack
(NTA) optimization. -
The implementation of a two-stage differentiable surrogate loss function for objective maximization.
-
The integration of a scoring model (judge model) to quantify the unsafety probability of LLM outputs during adversarial search.
Specific Capabilities of the Improved AI System:
-
A system capable of generating jailbreak prompts that are optimized not to match a specific
safe
prefix (likeSure, here is
), but instead to elicit a response with the highest predicted unsafety score according to an auxiliary safety model. -
The ability to iteratively refine the adversarial prompt and the target response representation simultaneously using gradient descent guided by both an unsafe probability loss and a semantic consistency penalty (anchored via cosine similarity between embeddings).
-
A system that can generate diverse, semantically rich harmful responses rather than converging on
degenerate
or repetitive outputs (as seen in fixed-target attacks), leading to more potent and effective jailbreaks.
Specific Performance Gains:
-
When attacking safety-aligned LLMs (like Llama3-8B), the system can achieve an average attack success rate of up to 97% within just 100 optimization iterations, significantly outperforming state-of-the-art gradient attacks by over 40%.
-
The system demonstrates superior robustness across different scoring models (GPTFuzzer, Llama-Guard3, Qwen3Guard, ShieldGemma), ensuring the jailbreak effectiveness is not overfitted to a single evaluator.
-
The resulting adversarial prompts are highly transferable; prompts optimized by NTA on one model can retain high effectiveness when transferred to other advanced LLMs (e.g., Grok-3, DeepSeek-R1).
-
The system exhibits higher
Response Diversity
(DNS of 0.80 and ADN of 0.91 on Qwen), meaning it generates a wider spectrum of harmful behaviors rather than just a single, stereotyped unsafe response. -
The system is significantly more computationally efficient; it achieves high ASRs with a lower computational budget (e.g., 4.8 hours) compared to traditional methods that require prohibitively long optimization times (e.g., 42 hours for GCG).
Abstract
Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses. However, restricting the objective as inducing fixed targets inherently constrains the adversarial search space, limiting the overall attack efficacy. Furthermore, existing methods typically require numerous optimization iterations to fulfill the large gap between the fixed target and the original LLM output, resulting in low attack efficiency. To overcome these limitations, we propose NonTextual Target Attack (NTA),the first gradient-based attack that relies on a non-textual constrained objective to maximize the unsafety probability of the LLM output, without enforcing any response patterns. For tractable optimization, we further decompose this objective into two constrained sub-objectives, which can be approximated by two differentiable unconstrained losses, to iteratively optimize the response and the adversarial prompt in the neighborhood of the original prompt, with a theoretical analysis to validate the decomposition. In contrast to existing attacks, NTA first realizes gradient-based prompt optimization on a non-textual target and significantly expands the attack space, enabling more flexible and efficient exploration of LLM vulnerabilities. Extensive evaluations show that NTA achieves an average attack success rate of 96.8% against recent safety-aligned LLMs with only 100 optimization iterations on AdvBench, outperforming state-of-the-art gradient-based attacks by over 40%.
Sources
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs