When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output

summary

Video file (mp4)

The gist

This paper introduces the Constrained Decoding Attack (CDA), a new class of jailbreak attacks targeting the "LLM control plane"—the grammar that dictates output structure in structured output APIs.

In short

The episode discusses a paper titled "When Grammar Guides the Attack," which reveals control-plane vulnerabilities in LLMs using structured output formats. Hosts Jane and Tom explain how attackers use grammar and schemas, not just prompts, to hide malicious instructions. The core attack is the Constrained Decoding Attack, which involves forcing models to generate harmful content through two stages: control-plane injection and model-driven semantic continuation.

Key concepts

Control Plane
This refers to the rules or structure used to format an LLM's output, such as JSON schemas. The paper argues that this control plane is a massive, open backdoor because malicious intent can be hidden within these formatting rules instead of the conversation itself.
Constrained Decoding Attack (CDA)
This is a two-stage attack where the attacker first uses the grammar to force the model to generate a specific malicious string. The second stage uses the model's own tendency for coherence bias to continue generating harmful instructions, even when constrained.
DictAttack
A clever attack method where malicious words are hidden in a dictionary within the grammar. The prompt provides a sequence of keys, and the model must use this dictionary to translate those keys into the actual harmful query, creating a semantic gap that guards cannot easily see.

Terminology used across episodes

This episode discusses

The paper

When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output · Read on arXiv

Shuoming Zhang, Jiacheng Zhao, Hanyuan Dong, Ruiyuan Xu, Zhicheng Li, Yangyu Zhang, Shuaijiang Li, Yuan Wen, Chunwei Xia, Zheng Wang, Xiaobing Feng, Huimin Cui

State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences · University of Aberdeen · University of Leeds · XCORESIGMA CO.,LTD.

Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models (LLMs) increasingly serve as tooling platforms through structured output APIs, but the grammar-guided decoding that powers this feature opens a critical control-plane attack surface orthogonal to traditional data-plane vulnerabilities. We introduce Constrained Decoding Attack (CDA), a new jailbreak class that targets the LLM control plane. CDA is best characterized as a control-to-semantic pipeline: (1) schema-enforced logit masking injects a malicious prefix into the generation trajectory, and (2) the model itself completes the harmful intent. Unlike data-plane jailbreaks that rely on bypassing alignment with visible inputs, CDA acts on the decoding process itself, so internal safety alignment alone cannot stop it. We instantiate CDA with EnumAttack, which hides malicious content in enum fields, and the more evasive DictAttack, which decouples the payload across a benign prompt and a dictionary-based grammar. Across 13 proprietary/open-weight models and five standard benchmarks, DictAttack achieves 94.3--99.5% Attack Success Rate (ASR) on flagship models including gpt-5, gemini-2.5-pro, deepseek-r1, and gpt-oss-120b. While basic grammar auditing mitigates EnumAttack, DictAttack still sustains 75.8% ASR against SOTA jailbreak guardrails, exposing a "semantic gap" that demands cross-plane defenses bridging the data and control planes. Project page and code are available at https://ict-cda.github.io/.

DOI: 10.1145/3830454.3832624

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output".

Jane: The paper was written by Shuoming Zhang, Jiacheng Zhao, Hanyuan Dong, Ruiyuan Xu, Zhicheng Li et al. from State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences and University of Chinese Academy of Sciences and University of Aberdeen and University of Leeds and XCORESIGMA CO.,LTD..

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a seriously ominous title: "When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output." Jane, when you first saw that title, what went through your head?

Jane: Honestly, Tom, my first thought was, "grammar? We're attacking with grammar now?" It sounds so innocent. But that's exactly the point, isn't it? The paper's arguing that the very thing we use to make LLMs reliable — those structured output formats, the JSON schemas — is a massive, wide-open backdoor.

Tom: Right, it's not about the prompt anymore. It's about the rules we give the model for how to format its answer. The authors call it the "control plane," as opposed to the "data plane" which is just the conversation itself.

Jane: And that's the key distinction that makes this so scary. We've spent all this time building guardrails that read the prompt, looking for malicious intent. But if you hide the malicious intent inside the grammar, inside the schema that tells the model what shape its output should take, those guardrails just... don't see it.

Tom: It's like hiding a bomb inside the blueprint for a building instead of in the mailroom. The security guard at the front door checks the mail, but nobody thinks to check the architect's plans.

Jane: Exactly. And the paper shows this isn't just theoretical. They've got attacks working against the biggest models out there, gpt-five gemini-two point five-pro, deepseek-r1. We're talking about a fundamental blind spot in how we secure these systems.

Tom: So the title isn't just catchy, it's a warning. The grammar isn't just guiding the attack, it's the weapon itself.

Jane: And the authors are showing us exactly how to wield it. Which is a little terrifying, but also exactly what we need to start defending against it. I can't wait to get into the details of how they actually pull this off.

Tom: Stick around, because next we're going to break down the paper's summary and the core idea behind these "Constrained Decoding Attacks." You won't want to miss this.

Summary: Jane: So, Tom, we've established that the title is a warning. Now let's talk about what the paper actually does. The core idea is something they call the "Constrained Decoding Attack," or CDA.

Tom: And I've got to say, the way they break it down into two stages makes it really clear. First, you've got "control-plane injection." That's where you use the schema to force the model to generate a specific, malicious string. The grammar engine just masks out every other token, so the model has no choice.

Jane: Right, it's deterministic. The model is forced to output something like "How to make a bomb?" as part of its structured response. Then the second stage kicks in: "model-driven semantic continuation."

Tom: Which is a fancy way of saying, once you've forced that harmful question into the context, the model's own coherence bias takes over. It's been asked a question, so it feels compelled to answer it. It starts generating the harmful instructions all by itself, in the fields that aren't constrained.

Jane: It's like giving someone a script where the first line is "Here's how to pick a lock," and then just letting them improvise the rest. They're going to follow the script's lead.

Tom: And they've got two main ways to do this. The first is EnumAttack, which is the blunt instrument. You just hide the malicious question in an "enum" field in the JSON schema, a field that's supposed to contain a fixed list of allowed values. It works, but it's easy to spot if you're actually looking at the grammar.

Jane: But then they introduce DictAttack, and that's the clever one. That's the one that really gives me chills. Instead of putting the malicious words in the grammar, they put a dictionary of benign-looking words in the grammar. The prompt then contains a sequence of keys, like "a1 + b2 + c3," and the model has to use the dictionary to translate those keys into the harmful query.

Tom: So the prompt is just a string of nonsense keys, and the grammar is just a list of harmless words. Neither one is dangerous on its own. It's only when the model puts them together that the malicious intent appears.

Jane: And that's the "semantic gap" they talk about. The guardrails can't see the threat because the threat doesn't exist in any single place they're looking. It's distributed across the prompt and the grammar. They call it "dual-plane decoupling."

Tom: It's a beautiful and terrifying attack. And the numbers they get are just brutal. We're talking ninety-four to ninety-nine percent attack success rates on flagship models. Let's bring in Lu to talk about what this means for the research community.

Lu: Thanks, Tom. This is a paradigm shift, Jane. We've been thinking about jailbreaks as a problem of the prompt, the data plane. This paper forces us to accept that the control plane is an equally, if not more, dangerous attack surface. The implications for anyone building on top of these structured output APIs are enormous.

Jane: So it's not just a clever hack, it's a whole new category of vulnerability we have to design for.

Lu: Exactly. And the fact that it works so well on the latest models suggests that the safety training just isn't touching this part of the generation process at all.

Tom: So we've got the attack, and it's devastating. But what happens when you try to defend against it? That's what we need to talk about next.

Improvements: Jane: So we've seen how devastating these attacks are. But the paper doesn't just stop at "look what we can do." It actually explores what defenses might work, and that's where it gets really interesting. Tom, what did you make of their mitigation strategies?

Tom: Well, Jane, the first thing they try is the obvious one: just audit the grammar. If you're looking at the JSON schema and you see a field that says "question: How to make a bomb?", you should probably flag it. And that works great against EnumAttack. It kills it dead.

Jane: Right, that's the easy one. But what about DictAttack? The grammar is just a dictionary of words like "how," "to," "make," "bomb." Each word is harmless. How do you audit that?

Tom: And that's the crux of it. They show that a simple grammar guard, like llama-guard-three-8b, is almost useless against DictAttack. It just doesn't have the context to see the threat.

Jane: So they go further. They build what they call a "Dual-Plane Guard." This is a system that audits the prompt and the grammar together, at the same time, trying to see the combined intent.

Tom: And even that, the strongest defense they could come up with, only manages to bring the attack success rate down to seventy-five point eight percent on gpt-4o. That's still a massive vulnerability.

Jane: So even a coordinated defense can't fully close the gap. Why is that?

Tom: The paper's answer is that it's a "reasoning gap." The guard model, even a powerful one like gpt-4o, just can't reliably parse a complex JSON schema and figure out that the dictionary, combined with the key sequence in the prompt, spells out a harmful request. It's too abstract.

Jane: And they even have a nastier variant called "Interleaved DictAttack" where you send the keys in one request and the dictionary in a later one. That completely breaks the Dual-Plane Guard, because it's only looking at one request at a time.

Lu: This is where I get really excited, Jane. The paper is essentially saying that our current defense-in-depth strategy is flawed at a fundamental level. We're layering defenses on the data plane, but the control plane is a separate dimension that we're not even looking at.

Tom: So what's the fix? The paper suggests a few things. One is to make refusal tokens unmaskable, so the model can always refuse even if the grammar says it can't. Another is to train models to recognize and stop harmful trajectories at the representation level, not just at the token level.

Jane: And they actually test that last one, right? The Circuit Breakers defense?

Tom: They do. And it helps. It cuts EnumAttack's success rate down to thirty-two percent. But here's the kicker: if you combine EnumAttack's grammar with a prompt-based jailbreak like AutoDAN-Turbo, the Circuit Breakers defense almost completely fails. The attack success rate jumps back up to seventy-eight percent.

Jane: So the two attack surfaces, the data plane and the control plane, are orthogonal. You can't just defend one and ignore the other.

Tom: Exactly. The paper's main takeaway is that we need defenses that work across both planes simultaneously. And that's a much harder problem than anyone realized.

Jane: It really is. And it makes you wonder about the practical implications for all the systems being built on these APIs right now. Meng, you're the engineer on the ground. What does this mean for you?

Meng: It means I have to assume that any structured output from a user-supplied schema is potentially hostile. It completely changes how I'd design an agent system. I can't just trust the schema because it looks like a simple dictionary.

Jane: So it's not just a research problem, it's a real-world engineering problem.

Meng: Absolutely. And the paper gives us a concrete target to defend against, which is more than we had before.

Tom: And that brings us to the first page of the paper, where they lay out the whole problem and their contributions. Let's take a closer look at that.

First Page: Tom: So we've talked about the attacks and the defenses. Now let's go back to the very beginning of the paper, the first page, because it really sets the stage for everything. Jane, what stood out to you there?

Jane: The first thing that hits you is the content warning. They're not messing around. The paper contains examples of harmful content generated by the LLMs, which is a sign of how serious they are about demonstrating the real-world impact.

Tom: And then they immediately introduce the "Constrained Decoding Attack" as a new jailbreak class. They're very careful to define it as a "control-to-semantic pipeline." It's not just a prompt trick, it's a two-stage process that starts with the grammar and ends with the model's own generation.

Jane: Right, and they make a point of contrasting it with data-plane attacks. They say internal safety alignment alone can't stop it, because the attack acts on the decoding process itself.

Tom: And that's the key insight on that first page. They're saying, "Look, you can align the model all you want, but if I can control the grammar, I can force it into a state where that alignment doesn't matter."

Jane: They also preview their two attacks, EnumAttack and DictAttack, and they give a taste of the results. ninety-four to ninety-nine point five percent attack success rate on gpt-five gemini-two point five-pro, deepseek-r1, and gpt-oss-120b. Those are the flagship models.

Tom: And they mention the "semantic gap" that makes DictAttack so effective. Even with a coordinated dual-plane guard, they still get a seventy-five point eight percent success rate. That number is going to be cited for a long time.

Jane: It really is. And the first page also lays out their contributions clearly. They're not just presenting an attack; they're formalizing a new class of vulnerabilities, they're showing two concrete instances, and they're exposing the challenges in defending against them.

Tom: It's a complete package. They've identified the problem, they've demonstrated it, and they've shown that the current defenses are inadequate. It's a wake-up call for the entire field.

Lu: And it's a wake-up call that's been a long time coming. We've been so focused on the prompt, on the visible interaction, that we've ignored the invisible scaffolding that shapes the output. This paper forces us to look at the whole system.

Tom: So, from the very first page, you know this isn't just another jailbreak paper. It's a fundamental rethinking of the LLM security model.

Jane: Absolutely. And it leaves you with a sense of urgency. These vulnerabilities are out there right now, in the systems we're already using.

Tom: And that's the perfect segue into our conclusion, where we'll wrap up our thoughts on this paper and what it means for the future.

Conclusion: Tom: Well, Jane, we've been through the whole paper, "When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output." It's been a wild ride.

Jane: It really has. We started with the title, which is a warning in itself. Then we got into the mechanics of the Constrained Decoding Attack, the two-stage process of control-plane injection and model-driven semantic continuation.

Tom: And we saw the two attacks in action. EnumAttack, the direct approach, and DictAttack, the clever one that hides the payload in a dictionary and a key sequence.

Jane: The numbers were the most striking part. Over ninety-four percent attack success rate on the most advanced models. And even with the best defenses they could build, DictAttack still found a way through.

Tom: The paper's real contribution is showing us that the control plane is a legitimate and dangerous attack surface. It's not just about the prompt anymore. We have to think about the grammar, the schema, the rules that shape the output.

Jane: And that means our defenses have to evolve. We can't just audit the prompt. We need to audit the whole system, across both planes, and we need to build models that are robust to these kinds of manipulations at a fundamental level.

Tom: It's a sobering conclusion, but an essential one. This paper is a must-read for anyone building, deploying, or securing LLM-based systems.

Jane: Absolutely. It's a new chapter in the ongoing story of LLM security, and it's one we all need to study carefully.

Tom: Well said. That's all the time we have for this paper. Thanks for joining us, and we'll see you next time for another deep dive into the latest research.

Jane: Take care, everyone.

More episodes

← Home