Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models
summary
The gist
The gist Unified autoregressive models enable multimodal backdoor attacks where a trigger can propagate malicious effects across multiple output modalities, making content more convincing and
In short
This work introduces Token by Token Backdoor Attack (ToBAC), a method that exploits unified autoregressive models to create multimodal backdoors. The attack uses 'autoregressive self-poisoning' where a trigger first generates a poisoned image, which then influences the model to produce malicious text. This allows triggers to work through both visual and textual inputs.
Key concepts
- Unified Autoregressive Models (UAMs)
- These are AI models that generate content, like images or text, one piece at a time in a sequence. The 'unified' aspect means the model handles different types of data—like images and text—within the same framework. This structure makes it vulnerable to attacks where a trigger can affect multiple outputs simultaneously.
- Token by Token Backdoor Attack (ToBAC)
- This is a specific attack that leverages the sequential nature of UAMs. The trigger causes the model to first create a poisoned image, and this image then feeds back into the text generation process. This 'autoregressive self-poisoning' ensures that a single textual trigger can reliably lead to both a malicious visual output and an accompanying harmful text.
- Black-box (Data Poisoning) Attack
- In this scenario, the attacker manipulates the model's training data by adding a small number of poisoned examples. These poisoned examples link an innocent text trigger to specific malicious image/text pairs. The goal is to teach the model a hidden association that allows it to generate harmful content when given the trigger later.
- Bidirectional Training Defense
- A proposed defense involves training the model on overlapping image-text pairs in both directions. This forces conflicting supervision signals during training, which disrupts the hidden 'trigger-target linkage' that the backdoor relies on. When trained this way, it significantly reduces the attack's success rate.
Terminology used across episodes
This episode discusses
- Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models · Paper Radio
- Analyzing The Language of Visual Tokens
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
- Erased but Not Forgotten: How Backdoors Compromise Concept Erasure · Paper Radio
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Diff-Cleanse: Identifying and Mitigating Backdoor Attacks in Diffusion Models
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
- Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Gemma: Open Models Based on Gemini Research and Technology
- Emu3: Next-Token Prediction is All You Need
- Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
- Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
- Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
The paper
Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models · Read on arXiv
TU Darmstadt & hessian.AI
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Token by Token, Compromised".
Elias: The gist Unified autoregressive models enable multimodal backdoor attacks where a trigger can propagate malicious effects across multiple output modalities,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So we’re looking at this paper, "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models." It tackles how these unified models, which handle text and vision together, can be tricked.
Elias: Yeah. The title itself points to the core issue: token by token poisoning where the trigger gets passed from one modality to another. It suggests that you can set up a link between something visual and something textual that then causes trouble later on.
Nadia: Exactly. What they’re looking at is how these unified models, which treat text and vision as one shared vocabulary, are vulnerable to this cross-modal attack chain.
Priya: From my side, I'm interested in what the data actually shows here—specifically how much of the model's output is affected when this backdoor is active.
Nadia: Right. And what they’re showing is that once you have that poisoned image token, it doesn't just stay an image; it feeds back into the text generation process to cause a malicious textual continuation.
Elias: It makes sense from a cryptographic angle because it shows how a single trigger can propagate across different layers of the model's architecture simultaneously.
Priya: So, if we’re talking about what this means for privacy and measurement research, are they showing that these attacks are subtle enough to stay hidden in the output quality?
Nadia: They show that the attack succeeds without losing visual quality, which is important because it means the user doesn't immediately notice something is wrong with what they see.
Elias: And they talk about two main ways this poisoning happens: black-box data poisoning and white-box model poisoning during fine-tuning. It’s interesting that they cover both approaches in this paper, "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models."
Priya: So, for the listener who doesn't code or study cryptography, what does that cross-modal chain actually translate to in terms of risk?
Nadia: It means someone could use a simple textual trigger to get a malicious image and then use that image to prompt the model into generating dangerous text.
Elias: And they detail the mechanism as "autoregressive self-poisoning," where the trigger first makes it generate a poisoned image, which then feeds back into the autoregressive context to elicit a malicious text token. It’s this chain of events that makes it work.
Priya: So, when we look at their findings on performance, they test this against models like LIQUID-7B and JANUSPRO. What are the actual numbers they report regarding attack success?
Nadia: For LIQUID, they report an ASRU of sixty-five point one zero percent for the smoking target scenario, and for JANUSPRO in the proud scenario, it hits ninety point three zero percent. Those are pretty high rates for this kind of coordinated output generation.
Elias: And they also demonstrate that because of how these unified models work, the distinction between a text-triggered attack and a vision-triggered attack is mostly about perspective rather than the actual mechanism used.
Priya: That’s interesting because it suggests the underlying vulnerability isn't tied to just one input modality, but how those modalities are linked within that single architecture.
Title and authors: Nadia: And they show an equivalence between text-triggered and vision-triggered attacks, meaning you can use an image externally to activate the same malicious behavior. That's a significant finding for understanding the unified model's weak points.
Elias: Now, looking at how they suggest fixing this, their defense strategy is enforcing bidirectional training on overlapping image-text pairs. They argue that alternating training directions disrupts that coherent trigger-target linkage by creating conflicting supervision signals.
Priya: So, from a privacy standpoint, what does this mean for defenses? Does forcing these bidirectional links actually help or just create a new kind of vulnerability we need to worry about?
Nadia: It’s presented as a way to disrupt that cross-modal coherence, and they show that when both directions are sampled with equal probability, the attack success rate drops dramatically. That suggests low poisoning rates might be enough on their own in the black-box setting.
Elias: But they also test some prompt-level security scanners like Prompt Guard two and LLM Guard PI, and those tools found they flagged virtually none of the prompts used in this paper <ref:2605.19227#pg1>.
Priya: And what about the limitation of these defenses? What stops them from being completely bypassable?
Nadia: The paper flags that prompt-level filtering alone likely won't be enough to stop everything, suggesting we need stronger data or model-level defenses instead.
Elias: They also pinpoint the "Aligned I2T link" as the most effective poisoning mechanism because it achieves high success rates while preserving stealth and utility better than other methods they tested.
Priya: So, to wrap up on the numbers, what’s the main conclusion we should take away from this study on "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models"?
Nadia: The main point is that unified models have a distinct vulnerability because they allow for backdoor attacks that span both image and text outputs through a single chain.
Elias: It confirms that the joint modeling capability introduces new security risks specifically because of cross-modal output consistency being exploitable.
Priya: And the practical implication is that defenses need to focus on disrupting the way modalities are linked during training rather than just filtering input prompts, as shown by the bidirectional training suggestion.
Nadia: So we’ve talked about how this attack works, what they found in terms of attack success rates on models like JANUSPRO and LIQUID-7B, and what defenses they propose against it. That’s a lot to chew on.
Elias: It really puts the focus on the structure of the unified model itself, showing that parameter sharing creates these new security pathways we have to account for when we build these systems.
Priya: I just think understanding this link between image tokens and text tokens is crucial because it shows how much more convincing fabricated content can become when those two outputs are manipulated together.
Nadia: Right, so the paper "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models" gives us a clear picture of how to exploit this unified generation capability.
The paper's summary: Nadia: So this paper shows how these unified models can have backdoors that jump between image and text outputs, it’s like setting up a chain where one thing triggers another later on.
Elias: It’s about "autoregressive self-poisoning," which means the trigger first makes the model generate a bad picture, and then that picture feeds back into the text generation to create a malicious sentence.
Priya: The key mechanism they use is this transitive chain, where the text trigger leads to a poisoned image, which then elicits a poisoned text response.
Nadia: And they explore two ways this happens: by poisoning the training data directly in a black-box way, or by embedding the trigger during fine-tuning in a white-box setting.
Elias: In that white-box scenario, they use a specific loss function to tie the image generation token to the text trigger while still letting it follow its normal autoregressive process.
Priya: What I’m seeing from the data is that even when they embed these backdoors, the overall utility of the model usually stays pretty stable, which is a big deal for real-world applications.
Nadia: They actually show that an aligned image-text link is the most effective poisoning method because it maintains high success rates without destroying how good the images or text actually are.
Elias: That’s interesting because it proves that the connection between what you see and what you read is where the real attack power lies in these unified architectures.
Priya: It also points out that this linkage means an external image can activate the same malicious behavior, which suggests any visual input is a potential trigger for text manipulation.
Nadia: And to defend against this, they suggest forcing the model to be trained on overlapping image-text pairs in both directions simultaneously to break that link.
Elias: That bidirectional training should create conflicting signals during training, which disrupts that coherent trigger-target association the attacker is trying to build.
Priya: It's a practical defense because when you sample those pairs equally, the attack success rate drops dramatically, showing that low poisoning rates are already pretty effective against this mechanism.
Nadia: But they also found that prompt-level scanners mostly miss these attacks, so we need stronger defenses focused on the data or model itself instead of just filtering the text instructions.
Elias: Right. It also confirmed that prompt filters don't catch these kinds of subtle triggers, which makes me think about what kind of structural changes we need to make in how models are trained to stop this propagation.
Priya: So, it seems the main implication for us is that the way we supervise these unified models needs to change fundamentally if we want them to be more trustworthy for multimodal content.
The paper's improvements: Tom: This paper isn't just about finding the vulnerability; they’re proposing ways to actually fix it, like using bidirectional training to stop that cross-modal link from forming in the first place.
Nadia: They suggest forcing the AI to learn associations between both poisoned pairs, so it sees both the trigger and the poisoned output together, which messes up that coherent connection.
Elias: That way of doing it should create conflicting signals during training, disrupting how those trigger-target associations are actually built in the model's brain.
Priya: It's a real mechanism because they showed that when you train both directions with equal weight, the attack success rate drops significantly.
Nadia: They also looked at tuning parameters; they found that setting a specific regularization value gives them the best balance—a strong attack rate while keeping the model’s standard output quality pretty stable.
Elias: That tuning point is interesting because it shows you can get a high success rate without immediately destroying how useful the AI actually is for general tasks.
Priya: The implication here for measurement research is that we need to focus on these structural training methods rather than just trying to clean the data after it's already been poisoned.
Nadia: They also pointed out that prompt-level scanners aren't doing much against this, so if you’re building defenses, you have to look at the model or the data level instead of just filtering what a user types in.
Elias: Exactly. The paper emphasizes that we need better internal detection mechanisms because surface-level filters won't catch these kinds of subtle cross-modal triggers.
Priya: So, it seems like the path forward is to redesign the training process itself to be more robust against these kinds of transitive poisoning chains.
Nadia: It’s a shift from just defending the input prompt to ensuring the entire model learns better ways to separate visual and textual information during its learning phase.
Conclusion: Tom: So we’re wrapping up on "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models" by summarizing how these unified models can have backdoors that jump between image and text outputs.
Nadia: Basically, they proved that this cross-modal chain—trigger to poisoned image to poisoned text—is real and exploitable across different model types.
Elias: It confirms that the way these models handle both modalities together creates a new type of security risk specifically because of how those outputs are linked.
Priya: What this means for measurement is that we have to be careful about what we measure, since a single trigger can mess with both the visual and textual results simultaneously.
Nadia: And the defense they landed on—bidirectional training—is a structural fix designed to break that link by introducing conflicting signals during the learning process.
Elias: That bidirectional approach is smart because it targets the coherence itself rather than just trying to patch a specific input, which is what prompt filters usually do.
Priya: It suggests that for future research on privacy and measurement, we need to look at how training data composition affects the model’s ability to maintain integrity across different modalities.
Nadia: Exactly. The paper shows us where the weaknesses are in these unified architectures, even when they seem pretty integrated.
Elias: It really puts the focus on how parameter sharing creates these new security pathways that we have to account for when we build these systems going forward.
More episodes
- 2610.10597-Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
- 2610.10608-From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
- 2610.10612-PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits