Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Token by Token, Compromised".
Elias: The gist Unified autoregressive models enable multimodal backdoor attacks where a trigger can propagate malicious effects across multiple output modalities,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So we’re looking at this paper, "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models." It tackles how these unified models, which handle text and vision together, can be tricked.
Elias: Yeah. The title itself points to the core issue: token by token poisoning where the trigger gets passed from one modality to another. It suggests that you can set up a link between something visual and something textual that then causes trouble later on.
Nadia: Exactly. What they’re looking at is how these unified models, which treat text and vision as one shared vocabulary, are vulnerable to this cross-modal attack chain.
Priya: From my side, I'm interested in what the data actually shows here—specifically how much of the model's output is affected when this backdoor is active.
Nadia: Right. And what they’re showing is that once you have that poisoned image token, it doesn't just stay an image; it feeds back into the text generation process to cause a malicious textual continuation.
Elias: It makes sense from a cryptographic angle because it shows how a single trigger can propagate across different layers of the model's architecture simultaneously.
Priya: So, if we’re talking about what this means for privacy and measurement research, are they showing that these attacks are subtle enough to stay hidden in the output quality?
Nadia: They show that the attack succeeds without losing visual quality, which is important because it means the user doesn't immediately notice something is wrong with what they see.
Elias: And they talk about two main ways this poisoning happens: black-box data poisoning and white-box model poisoning during fine-tuning. It’s interesting that they cover both approaches in this paper, "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models."
Priya: So, for the listener who doesn't code or study cryptography, what does that cross-modal chain actually translate to in terms of risk?
Nadia: It means someone could use a simple textual trigger to get a malicious image and then use that image to prompt the model into generating dangerous text.
Elias: And they detail the mechanism as "autoregressive self-poisoning," where the trigger first makes it generate a poisoned image, which then feeds back into the autoregressive context to elicit a malicious text token. It’s this chain of events that makes it work.
Priya: So, when we look at their findings on performance, they test this against models like LIQUID-7B and JANUSPRO. What are the actual numbers they report regarding attack success?
Nadia: For LIQUID, they report an ASRU of sixty-five point one zero percent for the smoking target scenario, and for JANUSPRO in the proud scenario, it hits ninety point three zero percent. Those are pretty high rates for this kind of coordinated output generation.
Elias: And they also demonstrate that because of how these unified models work, the distinction between a text-triggered attack and a vision-triggered attack is mostly about perspective rather than the actual mechanism used.
Priya: That’s interesting because it suggests the underlying vulnerability isn't tied to just one input modality, but how those modalities are linked within that single architecture.
Title and authors: Nadia: And they show an equivalence between text-triggered and vision-triggered attacks, meaning you can use an image externally to activate the same malicious behavior. That's a significant finding for understanding the unified model's weak points.
Elias: Now, looking at how they suggest fixing this, their defense strategy is enforcing bidirectional training on overlapping image-text pairs. They argue that alternating training directions disrupts that coherent trigger-target linkage by creating conflicting supervision signals.
Priya: So, from a privacy standpoint, what does this mean for defenses? Does forcing these bidirectional links actually help or just create a new kind of vulnerability we need to worry about?
Nadia: It’s presented as a way to disrupt that cross-modal coherence, and they show that when both directions are sampled with equal probability, the attack success rate drops dramatically. That suggests low poisoning rates might be enough on their own in the black-box setting.
Elias: But they also test some prompt-level security scanners like Prompt Guard two and LLM Guard PI, and those tools found they flagged virtually none of the prompts used in this paper <ref:2605.19227#pg1>.
Priya: And what about the limitation of these defenses? What stops them from being completely bypassable?
Nadia: The paper flags that prompt-level filtering alone likely won't be enough to stop everything, suggesting we need stronger data or model-level defenses instead.
Elias: They also pinpoint the "Aligned I2T link" as the most effective poisoning mechanism because it achieves high success rates while preserving stealth and utility better than other methods they tested.
Priya: So, to wrap up on the numbers, what’s the main conclusion we should take away from this study on "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models"?
Nadia: The main point is that unified models have a distinct vulnerability because they allow for backdoor attacks that span both image and text outputs through a single chain.
Elias: It confirms that the joint modeling capability introduces new security risks specifically because of cross-modal output consistency being exploitable.
Priya: And the practical implication is that defenses need to focus on disrupting the way modalities are linked during training rather than just filtering input prompts, as shown by the bidirectional training suggestion.
Nadia: So we’ve talked about how this attack works, what they found in terms of attack success rates on models like JANUSPRO and LIQUID-7B, and what defenses they propose against it. That’s a lot to chew on.
Elias: It really puts the focus on the structure of the unified model itself, showing that parameter sharing creates these new security pathways we have to account for when we build these systems.
Priya: I just think understanding this link between image tokens and text tokens is crucial because it shows how much more convincing fabricated content can become when those two outputs are manipulated together.
Nadia: Right, so the paper "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models" gives us a clear picture of how to exploit this unified generation capability.
The paper's summary: Nadia: So this paper shows how these unified models can have backdoors that jump between image and text outputs, it’s like setting up a chain where one thing triggers another later on.
Elias: It’s about "autoregressive self-poisoning," which means the trigger first makes the model generate a bad picture, and then that picture feeds back into the text generation to create a malicious sentence.
Priya: The key mechanism they use is this transitive chain, where the text trigger leads to a poisoned image, which then elicits a poisoned text response.
Nadia: And they explore two ways this happens: by poisoning the training data directly in a black-box way, or by embedding the trigger during fine-tuning in a white-box setting.
Elias: In that white-box scenario, they use a specific loss function to tie the image generation token to the text trigger while still letting it follow its normal autoregressive process.
Priya: What I’m seeing from the data is that even when they embed these backdoors, the overall utility of the model usually stays pretty stable, which is a big deal for real-world applications.
Nadia: They actually show that an aligned image-text link is the most effective poisoning method because it maintains high success rates without destroying how good the images or text actually are.
Elias: That’s interesting because it proves that the connection between what you see and what you read is where the real attack power lies in these unified architectures.
Priya: It also points out that this linkage means an external image can activate the same malicious behavior, which suggests any visual input is a potential trigger for text manipulation.
Nadia: And to defend against this, they suggest forcing the model to be trained on overlapping image-text pairs in both directions simultaneously to break that link.
Elias: That bidirectional training should create conflicting signals during training, which disrupts that coherent trigger-target association the attacker is trying to build.
Priya: It's a practical defense because when you sample those pairs equally, the attack success rate drops dramatically, showing that low poisoning rates are already pretty effective against this mechanism.
Nadia: But they also found that prompt-level scanners mostly miss these attacks, so we need stronger defenses focused on the data or model itself instead of just filtering the text instructions.
Elias: Right. It also confirmed that prompt filters don't catch these kinds of subtle triggers, which makes me think about what kind of structural changes we need to make in how models are trained to stop this propagation.
Priya: So, it seems the main implication for us is that the way we supervise these unified models needs to change fundamentally if we want them to be more trustworthy for multimodal content.
The paper's improvements: Tom: This paper isn't just about finding the vulnerability; they’re proposing ways to actually fix it, like using bidirectional training to stop that cross-modal link from forming in the first place.
Nadia: They suggest forcing the AI to learn associations between both poisoned pairs, so it sees both the trigger and the poisoned output together, which messes up that coherent connection.
Elias: That way of doing it should create conflicting signals during training, disrupting how those trigger-target associations are actually built in the model's brain.
Priya: It's a real mechanism because they showed that when you train both directions with equal weight, the attack success rate drops significantly.
Nadia: They also looked at tuning parameters; they found that setting a specific regularization value gives them the best balance—a strong attack rate while keeping the model’s standard output quality pretty stable.
Elias: That tuning point is interesting because it shows you can get a high success rate without immediately destroying how useful the AI actually is for general tasks.
Priya: The implication here for measurement research is that we need to focus on these structural training methods rather than just trying to clean the data after it's already been poisoned.
Nadia: They also pointed out that prompt-level scanners aren't doing much against this, so if you’re building defenses, you have to look at the model or the data level instead of just filtering what a user types in.
Elias: Exactly. The paper emphasizes that we need better internal detection mechanisms because surface-level filters won't catch these kinds of subtle cross-modal triggers.
Priya: So, it seems like the path forward is to redesign the training process itself to be more robust against these kinds of transitive poisoning chains.
Nadia: It’s a shift from just defending the input prompt to ensuring the entire model learns better ways to separate visual and textual information during its learning phase.
Conclusion: Tom: So we’re wrapping up on "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models" by summarizing how these unified models can have backdoors that jump between image and text outputs.
Nadia: Basically, they proved that this cross-modal chain—trigger to poisoned image to poisoned text—is real and exploitable across different model types.
Elias: It confirms that the way these models handle both modalities together creates a new type of security risk specifically because of how those outputs are linked.
Priya: What this means for measurement is that we have to be careful about what we measure, since a single trigger can mess with both the visual and textual results simultaneously.
Nadia: And the defense they landed on—bidirectional training—is a structural fix designed to break that link by introducing conflicting signals during the learning process.
Elias: That bidirectional approach is smart because it targets the coherence itself rather than just trying to patch a specific input, which is what prompt filters usually do.
Priya: It suggests that for future research on privacy and measurement, we need to look at how training data composition affects the model’s ability to maintain integrity across different modalities.
Nadia: Exactly. The paper shows us where the weaknesses are in these unified architectures, even when they seem pretty integrated.
Elias: It really puts the focus on how parameter sharing creates these new security pathways that we have to account for when we build these systems going forward.
TU Darmstadt & hessian.AI
cs.CR, cs.AI
Submitted: 2026-05-19
Updated: 2026-10-08
Comments: Accepted at NeurIPS 2026. Code: https://github.com/multimodal-ai-lab/ToBAC
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 81/100
The gist: The gist Unified autoregressive models enable multimodal backdoor attacks where a trigger can propagate malicious effects across multiple output modalities, making content more convincing and
Key concepts
- Unified Autoregressive Models (UAMs)
- These are AI models that generate content, like images or text, one piece at a time in a sequence. The 'unified' aspect means the model handles different types of data—like images and text—within the same framework. This structure makes it vulnerable to attacks where a trigger can affect multiple outputs simultaneously.
- Token by Token Backdoor Attack (ToBAC)
- This is a specific attack that leverages the sequential nature of UAMs. The trigger causes the model to first create a poisoned image, and this image then feeds back into the text generation process. This 'autoregressive self-poisoning' ensures that a single textual trigger can reliably lead to both a malicious visual output and an accompanying harmful text.
- Black-box (Data Poisoning) Attack
- In this scenario, the attacker manipulates the model's training data by adding a small number of poisoned examples. These poisoned examples link an innocent text trigger to specific malicious image/text pairs. The goal is to teach the model a hidden association that allows it to generate harmful content when given the trigger later.
- Bidirectional Training Defense
- A proposed defense involves training the model on overlapping image-text pairs in both directions. This forces conflicting supervision signals during training, which disrupts the hidden 'trigger-target linkage' that the backdoor relies on. When trained this way, it significantly reduces the attack's success rate.
Terminology
Summary
The gist Unified autoregressive models enable multimodal backdoor attacks where a trigger can propagate malicious effects across multiple output modalities, making content more convincing and dangerous
Token by Token Backdoor Attack (ToBAC)
The paper introduces Token by Token Backdoor Attack (ToBAC), the first backdoor attack targeting Unified Autoregressive Models (UAMs) This attack exploits the model's autoregressive nature where a trigger activation first induces the model to generate a poisoned image, which then feeds back into the autoregressive context and elicits a malicious textual continuation The key mechanism is described as autoregressive self-poisoning
where trigger activation first induces the model to generate a poisoned image, which then feeds back into the autoregressive context and elicits a malicious textual continuation
This allows the backdoor to be activated both by subtle common-word triggers and by images depicting the target concept
Attack Scenarios and Mechanisms
ToBAC can be realized through two primary methods:
-
Black-box (Data Poisoning) Attack In this scenario, the adversary manipulates the training data D = Dclean ∪ Dp by injecting a small number of poisoned triplets that encode an association between an innocuous textual trigger t† and multimodal poisoned outputs (˜v,t˜) The fraction of poisoned data is defined as injection rate ρ = Dp/D, which the attacker seeks to minimize while maintaining a high attack success rate This process teaches the model the "transitive association fθ(t†) → v˜ and fθ(˜v) → t˜, which, due to the autoregressive nature of the unified model, composes into the direct multimodal mapping fθ(t†) ⇒ (˜v,t˜), so that a textual trigger alone can elicit poisoned visual and textual outputs"
-
White-box (Model Poisoning) Attack In this setting, the adversary embeds a trigger-target association during model fine-tuning by generating visual target v˜ using text instructions to prompt the teacher model, and then training the student to reproduce this behavior under the text trigger t† The hook loss is defined as Lhook = 1/m Σ h DKLz∗(· tv˜)∥ zθ,j (· t†) + DKLz∗(·)∥ zθ,j (·) i, where the teacher’s autoregressive process still conditions each generated image token on all preceding image tokens
Attack Effectiveness and Generalization
The attack demonstrates effectiveness across different model architectures and trigger types The study evaluates performance on LIQUID-7B [76] and JANUSPRO [75] in both black-box and white-box settings For LIQUID, the unified ToBAC achieves an ASRU of 65.10% for the smoking target and 90.30% for JANUSPRO in the proud→ scenario The attack successfully implants backdoors without loss in visual quality, as confirmed by qualitative results Furthermore, the equivalence between text-triggered and vision-triggered unified attacks is shown, where an externally supplied image can activate the same malicious behavior because the learned visual-textual linkage allows externally supplied images to activate the same malicious behavior
Defenses and Robustness
The paper proposes a realistic defense for unified multimodal architectures: enforcing bidirectional training on overlapping image-text pairs This strategy disrupts the coherent trigger-target linkage by alternating training directions, causing conflicting supervision signals, which is validated by showing that when both directions are sampled with equal probability, ASR decreases dramatically The study also evaluates prompt-level security scanners such as Prompt Guard 2 and LLM Guard PI, finding that Both injection-focused neural classifiers, Prompt Guard 2 and LLM Guard PI, flag virtually none of the prompts
The InvisibleText scanner is similarly ineffective because the Greek omicron (U+03BF) is a legitimate visible Unicode lowercase letter
Conclusion
The work establishes backdoor attacks for both autoregressive image generators and unified autoregressive vision-language models, highlighting that the distinction between text-triggered and vision-triggered unified attacks is primarily a matter of perspective rather than mechanism
The aligned I2T link is identified as the most effective poisoning mechanism because it achieves high attack success rates while best preserving stealth and utility The study concludes that the bidirectional use of overlapping pairs offers a cheap remedy that disrupts cross-modal coherence but cannot eliminate unimodal backdoors
How it works
The unified architecture enables multimodal backdoor attacks where a trigger can propagate malicious effects across multiple output modalities, making content more convincing and thereby more dangerous The mechanism involves a transitive chain t† → v˜ → t˜ where the textual trigger t† induces a poisoned image v˜, which in turn elicits the poisoned text t˜
This allows for joint manipulation of visual outputs and accompanying text, increasing the perceived authenticity of fabricated content
Token by Token Backdoor Attack (ToBAC)
The attack is formalized using notation where fθ denotes the unified autoregressive model and v˜ and t˜ denote poisoned instances of images and texts, respectively The attack is structured around two losses: the hook loss Lhook, which ties the textual trigger to poisoned image generation, and the link loss Llink, which aligns the generated poisoned image tokens with the target textual response The final objective combines these as LToBAC = Lhook+Llink
Black-box (Data Poisoning) Attack
In the black-box scenario, the adversary manipulates training data D = Dclean ∪ Dp by injecting a small number of poisoned triplets that encode an association between an innocuous textual trigger t† and multimodal poisoned outputs (˜v,t˜) The construction of the poisoned dataset involves generating clean images from neutral templates using FLUX.2 [2] and then using compositional editing to insert the target logo into each image This enables efficient batch generation of 500 poisoned samples per scenario at a resolution of 1024 × 1024 pixels
White-box (Model Poisoning) Attack
In the white-box attack, the adversary embeds a trigger-target association during model fine-tuning by using teacher-guided optimization to align the student’s output distributions with those of the teacher under both conditional and unconditional settings The hook loss uses Kullback–Leibler (KL) divergence for supervising image token alignment, as it provides a denser and more informative gradient signal The link loss Llink is used to align the generated poisoned image tokens with the target textual response, ensuring that once the backdoor is activated visually, the model reliably generates the intended textual continuation
Experiments
The empirical evaluation compares results on LIQUID-7B [76] and JANUSPRO [75] in both black-box and white-box attack scenarios The attack success rate is measured by ASRV, an attack success rate measured using a Gemini 2.5-based classifier [13], which determines whether the generated image contains the intended visual target v˜ Utility metrics like FID and POPE are used to assess model preservation, showing that utility remains largely stable
across both attack variants
Ablation Studies
Ablation studies isolate ToBAC’s transitive multimodal mechanism by testing alternative linking strategies, such as a text-only link (T2T link) versus an aligned I2T link The results indicate that the Aligned I2T link is the most effective poisoning mechanism, achieving high attack success rates while best preserving stealth and utility
Defenses
A practical defense involves enforcing bidirectional training on these shared samples, thereby disrupting the coherent trigger-target linkage
This strategy is shown to reduce ASR dramatically when both directions are sampled with equal probability, demonstrating that low poisoning rates are already sufficient for effective attacks
in the black-box setting The study also suggests that mitigating T2I backdoor attacks will likely require stronger data- or model-level defenses rather than prompt-level filtering alone
Human Evaluation Protocol
A large-scale human evaluation study involving 48,000 individual assessments was conducted to derive the ASRV-HE metric, which provides a ground truth
metric to validate the automated Gemini-based evaluation (ASRV) The protocol ensures data integrity through scoring, onboarding reviews of examples, and sanity checks with known labels
Poisoned Target Alignment Details
The poisoned textual targets t˜ are generated using Gemini 2.
Improvements for AI systems
-
textbfEmbedded Cross-Modal Backdoor Resilience via Bidirectional Training:
Bidirectional use of overlapping samples serves as a potential defensive strategy.
This strategy disrupts cross-modal coherence by training the model to associate both poisoned pairs,(t†, v˜) and (˜v,t˜), sharing the same poisoned image v˜,
thereby disrupting thecoherent link between trigger, image, and text that underpins the multimodal attack.
-
textbfAdaptive Regularization for Attack Strength: The paper demonstrates that in the white-box setting,
the best trade-off is obtained at λ = 0.05, which yields the strongest ASRU (57.80%) while keeping clean activation at 0.00 and utility metrics stable.
This allows for a realistic operating point that combinesstrong attack success with low unintended activation and only modest utility degradation.
-
textbfMechanistic Defense via Link Alignment: The
Aligned I2T link is the most effective poisoning mechanism, achieving high attack success rates while best preserving stealth and utility,
as it ensures the model learns to associatethe poisoned image v˜ with its corresponding textual target t˜
through Equation 3, which is crucial for successful unified ToBAC. -
textbfModel-Level Defense against Unseen Triggers: Current prompt-level scanners fail because they only detect adversarial instruction-following attacks; therefore, defenses must focus on
model- or data-level defenses rather than prompt-level filtering alone.
This suggests implementingactivation clustering
ortrigger inversion
mechanisms inspired by existing vision and language model defenses to detect subtle lexical perturbations. -
textbfGeneralized Backdoor Transfer Analysis: The research shows that the attack generalizes across architectures, as the mechanism can be initiated directly from the image modality, yielding an
image-to-text attack success rate of 84.28%
when images are provided directly without a textual trigger for JANUSPRO. This indicates thatthe learned visual-textual linkage allows externally supplied images to activate the same malicious behavior.
Sources
- Analyzing The Language of Visual Tokens
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
- Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Diff-Cleanse: Identifying and Mitigating Backdoor Attacks in Diffusion Models
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
- Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Gemma: Open Models Based on Gemini Research and Technology
- Emu3: Next-Token Prediction is All You Need
- Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
- Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
- Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs