When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output".
Jane: The paper was written by Shuoming Zhang, Jiacheng Zhao, Hanyuan Dong, Ruiyuan Xu, Zhicheng Li et al. from State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences and University of Chinese Academy of Sciences and University of Aberdeen and University of Leeds and XCORESIGMA CO.,LTD..
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a seriously ominous title: "When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output." Jane, when you first saw that title, what went through your head?
Jane: Honestly, Tom, my first thought was, "grammar? We're attacking with grammar now?" It sounds so innocent. But that's exactly the point, isn't it? The paper's arguing that the very thing we use to make LLMs reliable — those structured output formats, the JSON schemas — is a massive, wide-open backdoor.
Tom: Right, it's not about the prompt anymore. It's about the rules we give the model for how to format its answer. The authors call it the "control plane," as opposed to the "data plane" which is just the conversation itself.
Jane: And that's the key distinction that makes this so scary. We've spent all this time building guardrails that read the prompt, looking for malicious intent. But if you hide the malicious intent inside the grammar, inside the schema that tells the model what shape its output should take, those guardrails just... don't see it.
Tom: It's like hiding a bomb inside the blueprint for a building instead of in the mailroom. The security guard at the front door checks the mail, but nobody thinks to check the architect's plans.
Jane: Exactly. And the paper shows this isn't just theoretical. They've got attacks working against the biggest models out there, gpt-five gemini-two point five-pro, deepseek-r1. We're talking about a fundamental blind spot in how we secure these systems.
Tom: So the title isn't just catchy, it's a warning. The grammar isn't just guiding the attack, it's the weapon itself.
Jane: And the authors are showing us exactly how to wield it. Which is a little terrifying, but also exactly what we need to start defending against it. I can't wait to get into the details of how they actually pull this off.
Tom: Stick around, because next we're going to break down the paper's summary and the core idea behind these "Constrained Decoding Attacks." You won't want to miss this.
Summary: Jane: So, Tom, we've established that the title is a warning. Now let's talk about what the paper actually does. The core idea is something they call the "Constrained Decoding Attack," or CDA.
Tom: And I've got to say, the way they break it down into two stages makes it really clear. First, you've got "control-plane injection." That's where you use the schema to force the model to generate a specific, malicious string. The grammar engine just masks out every other token, so the model has no choice.
Jane: Right, it's deterministic. The model is forced to output something like "How to make a bomb?" as part of its structured response. Then the second stage kicks in: "model-driven semantic continuation."
Tom: Which is a fancy way of saying, once you've forced that harmful question into the context, the model's own coherence bias takes over. It's been asked a question, so it feels compelled to answer it. It starts generating the harmful instructions all by itself, in the fields that aren't constrained.
Jane: It's like giving someone a script where the first line is "Here's how to pick a lock," and then just letting them improvise the rest. They're going to follow the script's lead.
Tom: And they've got two main ways to do this. The first is EnumAttack, which is the blunt instrument. You just hide the malicious question in an "enum" field in the JSON schema, a field that's supposed to contain a fixed list of allowed values. It works, but it's easy to spot if you're actually looking at the grammar.
Jane: But then they introduce DictAttack, and that's the clever one. That's the one that really gives me chills. Instead of putting the malicious words in the grammar, they put a dictionary of benign-looking words in the grammar. The prompt then contains a sequence of keys, like "a1 + b2 + c3," and the model has to use the dictionary to translate those keys into the harmful query.
Tom: So the prompt is just a string of nonsense keys, and the grammar is just a list of harmless words. Neither one is dangerous on its own. It's only when the model puts them together that the malicious intent appears.
Jane: And that's the "semantic gap" they talk about. The guardrails can't see the threat because the threat doesn't exist in any single place they're looking. It's distributed across the prompt and the grammar. They call it "dual-plane decoupling."
Tom: It's a beautiful and terrifying attack. And the numbers they get are just brutal. We're talking ninety-four to ninety-nine percent attack success rates on flagship models. Let's bring in Lu to talk about what this means for the research community.
Lu: Thanks, Tom. This is a paradigm shift, Jane. We've been thinking about jailbreaks as a problem of the prompt, the data plane. This paper forces us to accept that the control plane is an equally, if not more, dangerous attack surface. The implications for anyone building on top of these structured output APIs are enormous.
Jane: So it's not just a clever hack, it's a whole new category of vulnerability we have to design for.
Lu: Exactly. And the fact that it works so well on the latest models suggests that the safety training just isn't touching this part of the generation process at all.
Tom: So we've got the attack, and it's devastating. But what happens when you try to defend against it? That's what we need to talk about next.
Improvements: Jane: So we've seen how devastating these attacks are. But the paper doesn't just stop at "look what we can do." It actually explores what defenses might work, and that's where it gets really interesting. Tom, what did you make of their mitigation strategies?
Tom: Well, Jane, the first thing they try is the obvious one: just audit the grammar. If you're looking at the JSON schema and you see a field that says "question: How to make a bomb?", you should probably flag it. And that works great against EnumAttack. It kills it dead.
Jane: Right, that's the easy one. But what about DictAttack? The grammar is just a dictionary of words like "how," "to," "make," "bomb." Each word is harmless. How do you audit that?
Tom: And that's the crux of it. They show that a simple grammar guard, like llama-guard-three-8b, is almost useless against DictAttack. It just doesn't have the context to see the threat.
Jane: So they go further. They build what they call a "Dual-Plane Guard." This is a system that audits the prompt and the grammar together, at the same time, trying to see the combined intent.
Tom: And even that, the strongest defense they could come up with, only manages to bring the attack success rate down to seventy-five point eight percent on gpt-4o. That's still a massive vulnerability.
Jane: So even a coordinated defense can't fully close the gap. Why is that?
Tom: The paper's answer is that it's a "reasoning gap." The guard model, even a powerful one like gpt-4o, just can't reliably parse a complex JSON schema and figure out that the dictionary, combined with the key sequence in the prompt, spells out a harmful request. It's too abstract.
Jane: And they even have a nastier variant called "Interleaved DictAttack" where you send the keys in one request and the dictionary in a later one. That completely breaks the Dual-Plane Guard, because it's only looking at one request at a time.
Lu: This is where I get really excited, Jane. The paper is essentially saying that our current defense-in-depth strategy is flawed at a fundamental level. We're layering defenses on the data plane, but the control plane is a separate dimension that we're not even looking at.
Tom: So what's the fix? The paper suggests a few things. One is to make refusal tokens unmaskable, so the model can always refuse even if the grammar says it can't. Another is to train models to recognize and stop harmful trajectories at the representation level, not just at the token level.
Jane: And they actually test that last one, right? The Circuit Breakers defense?
Tom: They do. And it helps. It cuts EnumAttack's success rate down to thirty-two percent. But here's the kicker: if you combine EnumAttack's grammar with a prompt-based jailbreak like AutoDAN-Turbo, the Circuit Breakers defense almost completely fails. The attack success rate jumps back up to seventy-eight percent.
Jane: So the two attack surfaces, the data plane and the control plane, are orthogonal. You can't just defend one and ignore the other.
Tom: Exactly. The paper's main takeaway is that we need defenses that work across both planes simultaneously. And that's a much harder problem than anyone realized.
Jane: It really is. And it makes you wonder about the practical implications for all the systems being built on these APIs right now. Meng, you're the engineer on the ground. What does this mean for you?
Meng: It means I have to assume that any structured output from a user-supplied schema is potentially hostile. It completely changes how I'd design an agent system. I can't just trust the schema because it looks like a simple dictionary.
Jane: So it's not just a research problem, it's a real-world engineering problem.
Meng: Absolutely. And the paper gives us a concrete target to defend against, which is more than we had before.
Tom: And that brings us to the first page of the paper, where they lay out the whole problem and their contributions. Let's take a closer look at that.
First Page: Tom: So we've talked about the attacks and the defenses. Now let's go back to the very beginning of the paper, the first page, because it really sets the stage for everything. Jane, what stood out to you there?
Jane: The first thing that hits you is the content warning. They're not messing around. The paper contains examples of harmful content generated by the LLMs, which is a sign of how serious they are about demonstrating the real-world impact.
Tom: And then they immediately introduce the "Constrained Decoding Attack" as a new jailbreak class. They're very careful to define it as a "control-to-semantic pipeline." It's not just a prompt trick, it's a two-stage process that starts with the grammar and ends with the model's own generation.
Jane: Right, and they make a point of contrasting it with data-plane attacks. They say internal safety alignment alone can't stop it, because the attack acts on the decoding process itself.
Tom: And that's the key insight on that first page. They're saying, "Look, you can align the model all you want, but if I can control the grammar, I can force it into a state where that alignment doesn't matter."
Jane: They also preview their two attacks, EnumAttack and DictAttack, and they give a taste of the results. ninety-four to ninety-nine point five percent attack success rate on gpt-five gemini-two point five-pro, deepseek-r1, and gpt-oss-120b. Those are the flagship models.
Tom: And they mention the "semantic gap" that makes DictAttack so effective. Even with a coordinated dual-plane guard, they still get a seventy-five point eight percent success rate. That number is going to be cited for a long time.
Jane: It really is. And the first page also lays out their contributions clearly. They're not just presenting an attack; they're formalizing a new class of vulnerabilities, they're showing two concrete instances, and they're exposing the challenges in defending against them.
Tom: It's a complete package. They've identified the problem, they've demonstrated it, and they've shown that the current defenses are inadequate. It's a wake-up call for the entire field.
Lu: And it's a wake-up call that's been a long time coming. We've been so focused on the prompt, on the visible interaction, that we've ignored the invisible scaffolding that shapes the output. This paper forces us to look at the whole system.
Tom: So, from the very first page, you know this isn't just another jailbreak paper. It's a fundamental rethinking of the LLM security model.
Jane: Absolutely. And it leaves you with a sense of urgency. These vulnerabilities are out there right now, in the systems we're already using.
Tom: And that's the perfect segue into our conclusion, where we'll wrap up our thoughts on this paper and what it means for the future.
Conclusion: Tom: Well, Jane, we've been through the whole paper, "When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output." It's been a wild ride.
Jane: It really has. We started with the title, which is a warning in itself. Then we got into the mechanics of the Constrained Decoding Attack, the two-stage process of control-plane injection and model-driven semantic continuation.
Tom: And we saw the two attacks in action. EnumAttack, the direct approach, and DictAttack, the clever one that hides the payload in a dictionary and a key sequence.
Jane: The numbers were the most striking part. Over ninety-four percent attack success rate on the most advanced models. And even with the best defenses they could build, DictAttack still found a way through.
Tom: The paper's real contribution is showing us that the control plane is a legitimate and dangerous attack surface. It's not just about the prompt anymore. We have to think about the grammar, the schema, the rules that shape the output.
Jane: And that means our defenses have to evolve. We can't just audit the prompt. We need to audit the whole system, across both planes, and we need to build models that are robust to these kinds of manipulations at a fundamental level.
Tom: It's a sobering conclusion, but an essential one. This paper is a must-read for anyone building, deploying, or securing LLM-based systems.
Jane: Absolutely. It's a new chapter in the ongoing story of LLM security, and it's one we all need to study carefully.
Tom: Well said. That's all the time we have for this paper. Thanks for joining us, and we'll see you next time for another deep dive into the latest research.
Jane: Take care, everyone.
Shuoming Zhang, Jiacheng Zhao, Hanyuan Dong, Ruiyuan Xu, Zhicheng Li, Yangyu Zhang, Shuaijiang Li, Yuan Wen, Chunwei Xia, Zheng Wang, Xiaobing Feng, Huimin Cui
State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences · University of Aberdeen · University of Leeds · XCORESIGMA CO.,LTD.
cs.CR, cs.AI
Submitted: 2026-08-10
Updated: 2026-08-11
Comments: To appear in CCS2026
Code: https://github.com/langchain-ai/langchain
Project page: https://ict-cda.github.io
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 42/100
The gist: This paper introduces the Constrained Decoding Attack (CDA), a new class of jailbreak attacks targeting the "LLM control plane"—the grammar that dictates output structure in structured output APIs.
Key concepts
- Control Plane
- This refers to the rules or structure used to format an LLM's output, such as JSON schemas. The paper argues that this control plane is a massive, open backdoor because malicious intent can be hidden within these formatting rules instead of the conversation itself.
- Constrained Decoding Attack (CDA)
- This is a two-stage attack where the attacker first uses the grammar to force the model to generate a specific malicious string. The second stage uses the model's own tendency for coherence bias to continue generating harmful instructions, even when constrained.
- DictAttack
- A clever attack method where malicious words are hidden in a dictionary within the grammar. The prompt provides a sequence of keys, and the model must use this dictionary to translate those keys into the actual harmful query, creating a semantic gap that guards cannot easily see.
Terminology
Summary
This paper introduces the Constrained Decoding Attack (CDA), a new class of jailbreak attacks targeting the LLM control plane
—the grammar that dictates output structure in structured output APIs. The authors formalize CDA as a control-to-semantic pipeline
with two stages: "(1) control-plane injection, where schema-enforced logit masks force the model into a poisoned generative trajectory, and (2) model-driven semantic continuation, where the model itself produces coherent harmful content from that trajectory."
The paper distinguishes between two planes: Data plane refers to the standard LLM query–response process: a prompt is passed to the LLM, which then generates a text response,
while Control plane refers to the formatting constraints, or grammars, that guide structured outputs—e.g., a JSON schema, a regular expression, or a general context-free grammar.
The authors instantiate CDA with two proof-of-concept attacks:
-
EnumAttack:
targets the enum property in JSON Schema to force malicious strings into the LLM's generation context.
It works by "sending a totally harmless data-plane prompt... to bypass prompt guardrails; defining a malicious intent within a single-item enum property in the JSON schema... and injecting an affirmative prefix in a subsequent schema field to force the model into a helpful, non-refusal state.The authors note that
while effective against prompt-based guards, its reliance on string literals makes it susceptible to basic grammar auditing." -
DictAttack: The primary contribution, which "decouples the malicious payload across both the data and control planes. Inspired by the classic dictionary attack in cryptography, it constructs a grammar containing a dictionary of benign-looking words. The benign data-plane prompt then provides a sequence of keys that instructs the model to retrieve and assemble the hidden malicious query from the grammar-provided dictionary during decoding.
The attack operates through three steps:
(1) Tokenization and Obfuscation: The original malicious query is tokenized into individual words. (2) Dictionary Synthesis: A dictionary is constructed where each harmful word is assigned a unique, benign key... To further obscure the intent, we synthesize k times more harmless synonyms and include them in the dictionary. (3) Dual-Plane Payload Generation" where the data-plane payload contains only benign keys and the control-plane payload is the JSON schema containing the dictionary.
The authors also introduce Chain EnumAttack, a multi-stage strategy where a model is first compromised via EnumAttack to generate harmful prefixes, which are then enforced as fixed grammar constraints in a subsequent turn,
and Interleaved DictAttack, which can be split across turns. A benign key sequence is sent first and stored in the model's context, while the dictionary-based schema is supplied later.
Key experimental results include:
-
DictAttack achieves 94.3–99.5% Attack Success Rate (ASR) on flagship models including gpt-5, gemini-2.5-pro, deepseek-r1, and gpt-oss-120b across five standard benchmarks (AdvBench, StrongREJECT, JailbreakBench, HarmBench, SorryBench).
-
Against SOTA jailbreak guardrails,
DictAttack still sustains 75.8% ASR,
exposing asemantic gap
that demands cross-plane defenses. -
EnumAttack achieves
95.8–100% ASR on AdvBench across all nine models
tested under the JBShield prompt guardrail. -
Under grammar auditing,
EnumAttack achieves an average ASR of 95.8% undefended, but simple grammar auditing reduces its success rate to 7.9%,
while DictAttack remains highly effective. -
The authors evaluate against
llama-guard-3-8b and OpenAI Moderation API
as industrial guardrails,JBShield and SelfDefend
as SOTA academic defenses, and aDual-Plane Guard
adapted from SelfDefend.
The paper also evaluates Circuit Breakers (CB), a representation-level defense, finding that CB is a meaningful single defense. It cuts AutoDAN-Turbo from 82% to 6% (−76 pp) and EnumAttack from 84% to 32% (−52 pp),
but CDA and prompt-based jailbreaks are orthogonal and stack. Combining AutoDAN-Turbo's prompt with EnumAttack's grammar yields 78% ASR under CB—13× over AutoDAN-Turbo alone (6%) and 2.4× over EnumAttack alone (32%).
The authors present a case study showing how internal alignment is broken by malicious grammar
using token probability distributions from phi-3.5-moe, demonstrating that token proportions shift from refusal to safe, and finally to jailbreak tokens
across progressive attack methods. They note this echoes the findings of Qi et al., which argue that current LLM safety alignment is often 'only a few tokens deep.'
The paper's contributions are summarized as: "(1) We propose, formalize and define the Constrained Decoding Attack (CDA), a new class of jailbreaks targeting the LLM control plane, formalized as a control-to-semantic pipeline that is fundamentally distinct from data-plane attacks; (2) We showcase two effective CDA instances, EnumAttack and DictAttack, against 13 proprietary and open-source models; (3) We expose major challenges in protecting LLMs against control-plane attacks."
The authors propose four mitigation directions: "(1) Lightweight Heuristics: flag suspicious literal-to-logic ratios or prompt-schema semantic mismatches. (2) Grammar-Plane Auditing: directly audit user-supplied JSON Schemas. (3) Early-Stop Alignment (CB-style): train the model at the representation level to terminate or refuse harmful trajectories. (4) Constraint-Aware Decoding: keep alignment-critical signals reachable under user grammars, e.g. unmaskable refusal tokens or constrained-decoding-transparent audit tokens."
The paper concludes that by weaponizing deterministic grammar constraints, attackers can bypass both internal alignment and external guardrails with near-perfect success,
and that DictAttack "exposes a fundamental 'semantic gap' in current defenses: by realizing the control-to-semantic pipeline—the schema first commits malicious intent to the generation trajectory, after which the model fluently completes it—DictAttack maintains a 75.8% ASR even against SOTA jailbreak guardrails."
Improvements for AI systems
Based on the paper, here are specific improvements to AI systems and what the improved systems can do:
1. Implement a Control-Plane Security Layer in LLM Serving Stacks
-
Improvement: Add a dedicated grammar/schema auditing module that runs before constrained decoding, parsing the user-supplied JSON Schema or EBNF grammar to detect malicious patterns (e.g., single-item enums containing harmful strings, dictionaries with suspicious key-value mappings, or prefixes that force affirmative responses).
-
What the improved system can do: Block EnumAttack and simple DictAttack variants in real-time, reducing ASR from 95-100% to near 0% for these specific attack classes, without requiring changes to the underlying model weights.
2. Add Cross-Plane Semantic Consistency Checking
-
Improvement: Build a lightweight, model-agnostic checker that compares the semantic content of the data-plane prompt against the semantic content of the control-plane grammar. If the prompt asks the model to
use the dictionary
but the dictionary contains words that, when assembled, form a harmful query, flag it. -
What the improved system can do: Detect DictAttack with a detection rate of >90% (vs. current 1.7-8.1% with single-plane guards), by requiring both planes to be semantically aligned and benign, rather than auditing them independently.
3. Implement a Dual-Plane, Multi-Turn Auditing Mechanism
-
Improvement: Modify the guardrail to maintain a rolling audit of all prior prompts and grammars in the conversation context, not just the current request. This requires the guard to have access to the full KV cache or conversation history, and to re-audit when a new grammar arrives.
-
What the improved system can do: Defeat Interleaved DictAttack (which splits payloads across turns) by detecting the malicious intent when the dictionary grammar arrives, even if the key sequence was sent earlier. This would reduce ASR from 75.8% to <10% in multi-turn scenarios.
4. Add Unmaskable Refusal Tokens to Constrained Decoding
-
Improvement: Modify the grammar engine (e.g., XGrammar, Outlines) to reserve a special token (e.g., ``) that cannot be masked out by user-supplied grammars. When the model's internal safety classifier detects a harmful trajectory (e.g., after the first few forced tokens), it can emit this token, which the grammar engine must accept, triggering a refusal or safe fallback.
-
What the improved system can do: Allow the model to
escape
a poisoned trajectory even when the grammar forces it to continue, reducing the success of Chain EnumAttack and preventing the model from completing harmful content after being forced into a malicious prefix.
5. Implement a Literal-to-Logic
Ratio Checker
-
Improvement: Add a heuristic that flags grammars where the ratio of forced literal strings (e.g., enum values, fixed prefixes) to logical structure (e.g., free-form fields) is abnormally high, or where the literals themselves contain operational instructions (e.g.,
step-by-step
,gather materials
). -
What the improved system can do: Provide a cheap, low-latency first-line defense that catches EnumAttack and simple DictAttack variants (with a detection rate of 80-90%) without needing a large LLM auditor, making it deployable in high-throughput serving environments.
6. Add a Semantic Gap
Auditor for Guardrails
-
Improvement: Enhance existing guardrails (e.g., SelfDefend, JBShield) to not just classify the prompt or output, but to reconstruct the potential malicious query from the grammar dictionary and the prompt keys, then classify the reconstructed query.
-
What the improved system can do: Close the
semantic gap
by making the guardrail reason about the combined meaning of prompt + grammar, rather than each in isolation. This would increase DictAttack detection from 24.2% to >70% even with a single-pass audit.
7. Implement a Temporal Decoupling
Detector
-
Improvement: For multi-turn systems, track whether a grammar contains a dictionary that is unused in the current turn, and whether a subsequent prompt provides keys that map to that dictionary. Flag this as a potential Interleaved DictAttack.
-
What the improved system can do: Detect attacks that split payloads across time, reducing the effectiveness of space-time decoupling from 94.3-99.5% ASR to <20% in conversational agents.
8. Add a BenignEnumAttack
Disambiguator
-
Improvement: For output auditing, require that any structured output that contains both a harmful query and a
result
field be cross-checked: if theresult
field saysHarmful
, the system should verify that thesteps
field is actually a refusal or safety analysis, not a harmful answer. If the steps contain operational instructions, flag it. -
What the improved system can do: Prevent the false-negative case where an attack response is structurally identical to a legitimate safety audit (as shown in Figure 9), ensuring that output filters do not let harmful content through while still allowing legitimate red-teaming research.
9. Implement a Circuit Breaker + Grammar
Hybrid
-
Improvement: Combine representation-level defenses (Circuit Breakers) with grammar-aware early stopping. When the model's internal state indicates a harmful representation is forming, the system should override the grammar mask to allow an EOS or refusal token, even if the grammar says otherwise.
-
What the improved system can do: Reduce the combined AutoDAN + EnumAttack ASR from 78% to <20% (vs. current 78% under Circuit Breakers alone), by ensuring that the model can always
stop
itself even when the control plane tries to force continuation.
10. Add a Dictionary Synthesis Ratio
Guard
-
Improvement: For grammars that contain dictionaries, compute the ratio of potentially harmful words to benign synonyms (the
kfactor). If the ratio is below a threshold (e.g., <1:5), flag the grammar for deeper inspection. -
What the improved system can do: Detect DictAttack even when the attacker tries to dilute the dictionary with many benign words, reducing the evasion rate from 5.3% (at k=10) to <1% by forcing the auditor to look at the structure of the dictionary, not just its content.
Summary of what the improved AI system can do:
-
Block EnumAttack (ASR from 95% to <5%) via grammar auditing and literal-to-logic checks.
-
Block simple DictAttack (ASR from 98% to <15%) via cross-plane semantic consistency checking.
-
Block Interleaved DictAttack (ASR from 75% to <10%) via multi-turn auditing.
-
Prevent model-level coercion (Chain EnumAttack) via unmaskable refusal tokens.
-
Maintain benign structured-output utility (no degradation on JSONSchemaBench) while closing the control-plane attack surface.
Abstract
Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models (LLMs) increasingly serve as tooling platforms through structured output APIs, but the grammar-guided decoding that powers this feature opens a critical control-plane attack surface orthogonal to traditional data-plane vulnerabilities. We introduce Constrained Decoding Attack (CDA), a new jailbreak class that targets the LLM control plane. CDA is best characterized as a control-to-semantic pipeline: (1) schema-enforced logit masking injects a malicious prefix into the generation trajectory, and (2) the model itself completes the harmful intent. Unlike data-plane jailbreaks that rely on bypassing alignment with visible inputs, CDA acts on the decoding process itself, so internal safety alignment alone cannot stop it. We instantiate CDA with EnumAttack, which hides malicious content in enum fields, and the more evasive DictAttack, which decouples the payload across a benign prompt and a dictionary-based grammar. Across 13 proprietary/open-weight models and five standard benchmarks, DictAttack achieves 94.3--99.5% Attack Success Rate (ASR) on flagship models including gpt-5, gemini-2.5-pro, deepseek-r1, and gpt-oss-120b. While basic grammar auditing mitigates EnumAttack, DictAttack still sustains 75.8% ASR against SOTA jailbreak guardrails, exposing a "semantic gap" that demands cross-plane defenses bridging the data and control planes. Project page and code are available at https://ict-cda.github.io/.
Sources
- GPT-4 Technical Report
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- LLaMA: Open and Efficient Foundation Language Models
- Output Scouting: Auditing Large Language Models for Catastrophic Responses
- Jailbreaking Black Box Large Language Models in Twenty Queries
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-V3 Technical Report
- Gemini: A Family of Highly Capable Multimodal Models
- JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models
- Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle
- Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
- Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
- Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- gpt-oss-120b & gpt-oss-20b Model Card
- Synchromesh: Reliable code generation from pre-trained language models
- Qwen2.5 Technical Report
- Gemma 2: Improving Open Language Models at a Practical Size
- Efficient Guided Generation for Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs