Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models

arXiv:2605.19227 · cs.CR, cs.AI · Submitted 2026-05-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Token by Token, Compromised".

Elias: The gist Unified autoregressive models enable multimodal backdoor attacks where a trigger can propagate malicious effects across multiple output modalities,

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So we’re looking at this paper, "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models." It tackles how these unified models, which handle text and vision together, can be tricked.

Elias: Yeah. The title itself points to the core issue: token by token poisoning where the trigger gets passed from one modality to another. It suggests that you can set up a link between something visual and something textual that then causes trouble later on.

Nadia: Exactly. What they’re looking at is how these unified models, which treat text and vision as one shared vocabulary, are vulnerable to this cross-modal attack chain.

Priya: From my side, I'm interested in what the data actually shows here—specifically how much of the model's output is affected when this backdoor is active.

Nadia: Right. And what they’re showing is that once you have that poisoned image token, it doesn't just stay an image; it feeds back into the text generation process to cause a malicious textual continuation.

Elias: It makes sense from a cryptographic angle because it shows how a single trigger can propagate across different layers of the model's architecture simultaneously.

Priya: So, if we’re talking about what this means for privacy and measurement research, are they showing that these attacks are subtle enough to stay hidden in the output quality?

Nadia: They show that the attack succeeds without losing visual quality, which is important because it means the user doesn't immediately notice something is wrong with what they see.

Elias: And they talk about two main ways this poisoning happens: black-box data poisoning and white-box model poisoning during fine-tuning. It’s interesting that they cover both approaches in this paper, "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models."

Priya: So, for the listener who doesn't code or study cryptography, what does that cross-modal chain actually translate to in terms of risk?

Nadia: It means someone could use a simple textual trigger to get a malicious image and then use that image to prompt the model into generating dangerous text.

Elias: And they detail the mechanism as "autoregressive self-poisoning," where the trigger first makes it generate a poisoned image, which then feeds back into the autoregressive context to elicit a malicious text token. It’s this chain of events that makes it work.

Priya: So, when we look at their findings on performance, they test this against models like LIQUID-7B and JANUSPRO. What are the actual numbers they report regarding attack success?

Nadia: For LIQUID, they report an ASRU of sixty-five point one zero percent for the smoking target scenario, and for JANUSPRO in the proud scenario, it hits ninety point three zero percent. Those are pretty high rates for this kind of coordinated output generation.

Elias: And they also demonstrate that because of how these unified models work, the distinction between a text-triggered attack and a vision-triggered attack is mostly about perspective rather than the actual mechanism used.

Priya: That’s interesting because it suggests the underlying vulnerability isn't tied to just one input modality, but how those modalities are linked within that single architecture.

Title and authors: Nadia: And they show an equivalence between text-triggered and vision-triggered attacks, meaning you can use an image externally to activate the same malicious behavior. That's a significant finding for understanding the unified model's weak points.

Elias: Now, looking at how they suggest fixing this, their defense strategy is enforcing bidirectional training on overlapping image-text pairs. They argue that alternating training directions disrupts that coherent trigger-target linkage by creating conflicting supervision signals.

Priya: So, from a privacy standpoint, what does this mean for defenses? Does forcing these bidirectional links actually help or just create a new kind of vulnerability we need to worry about?

Nadia: It’s presented as a way to disrupt that cross-modal coherence, and they show that when both directions are sampled with equal probability, the attack success rate drops dramatically. That suggests low poisoning rates might be enough on their own in the black-box setting.

Elias: But they also test some prompt-level security scanners like Prompt Guard two and LLM Guard PI, and those tools found they flagged virtually none of the prompts used in this paper <ref:2605.19227#pg1>.

Priya: And what about the limitation of these defenses? What stops them from being completely bypassable?

Nadia: The paper flags that prompt-level filtering alone likely won't be enough to stop everything, suggesting we need stronger data or model-level defenses instead.

Elias: They also pinpoint the "Aligned I2T link" as the most effective poisoning mechanism because it achieves high success rates while preserving stealth and utility better than other methods they tested.

Priya: So, to wrap up on the numbers, what’s the main conclusion we should take away from this study on "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models"?

Nadia: The main point is that unified models have a distinct vulnerability because they allow for backdoor attacks that span both image and text outputs through a single chain.

Elias: It confirms that the joint modeling capability introduces new security risks specifically because of cross-modal output consistency being exploitable.

Priya: And the practical implication is that defenses need to focus on disrupting the way modalities are linked during training rather than just filtering input prompts, as shown by the bidirectional training suggestion.

Nadia: So we’ve talked about how this attack works, what they found in terms of attack success rates on models like JANUSPRO and LIQUID-7B, and what defenses they propose against it. That’s a lot to chew on.

Elias: It really puts the focus on the structure of the unified model itself, showing that parameter sharing creates these new security pathways we have to account for when we build these systems.

Priya: I just think understanding this link between image tokens and text tokens is crucial because it shows how much more convincing fabricated content can become when those two outputs are manipulated together.

Nadia: Right, so the paper "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models" gives us a clear picture of how to exploit this unified generation capability.

The paper's summary: Nadia: So this paper shows how these unified models can have backdoors that jump between image and text outputs, it’s like setting up a chain where one thing triggers another later on.

Elias: It’s about "autoregressive self-poisoning," which means the trigger first makes the model generate a bad picture, and then that picture feeds back into the text generation to create a malicious sentence.

Priya: The key mechanism they use is this transitive chain, where the text trigger leads to a poisoned image, which then elicits a poisoned text response.

Nadia: And they explore two ways this happens: by poisoning the training data directly in a black-box way, or by embedding the trigger during fine-tuning in a white-box setting.

Elias: In that white-box scenario, they use a specific loss function to tie the image generation token to the text trigger while still letting it follow its normal autoregressive process.

Priya: What I’m seeing from the data is that even when they embed these backdoors, the overall utility of the model usually stays pretty stable, which is a big deal for real-world applications.

Nadia: They actually show that an aligned image-text link is the most effective poisoning method because it maintains high success rates without destroying how good the images or text actually are.

Elias: That’s interesting because it proves that the connection between what you see and what you read is where the real attack power lies in these unified architectures.

Priya: It also points out that this linkage means an external image can activate the same malicious behavior, which suggests any visual input is a potential trigger for text manipulation.

Nadia: And to defend against this, they suggest forcing the model to be trained on overlapping image-text pairs in both directions simultaneously to break that link.

Elias: That bidirectional training should create conflicting signals during training, which disrupts that coherent trigger-target association the attacker is trying to build.

Priya: It's a practical defense because when you sample those pairs equally, the attack success rate drops dramatically, showing that low poisoning rates are already pretty effective against this mechanism.

Nadia: But they also found that prompt-level scanners mostly miss these attacks, so we need stronger defenses focused on the data or model itself instead of just filtering the text instructions.

Elias: Right. It also confirmed that prompt filters don't catch these kinds of subtle triggers, which makes me think about what kind of structural changes we need to make in how models are trained to stop this propagation.

Priya: So, it seems the main implication for us is that the way we supervise these unified models needs to change fundamentally if we want them to be more trustworthy for multimodal content.

The paper's improvements: Tom: This paper isn't just about finding the vulnerability; they’re proposing ways to actually fix it, like using bidirectional training to stop that cross-modal link from forming in the first place.

Nadia: They suggest forcing the AI to learn associations between both poisoned pairs, so it sees both the trigger and the poisoned output together, which messes up that coherent connection.

Elias: That way of doing it should create conflicting signals during training, disrupting how those trigger-target associations are actually built in the model's brain.

Priya: It's a real mechanism because they showed that when you train both directions with equal weight, the attack success rate drops significantly.

Nadia: They also looked at tuning parameters; they found that setting a specific regularization value gives them the best balance—a strong attack rate while keeping the model’s standard output quality pretty stable.

Elias: That tuning point is interesting because it shows you can get a high success rate without immediately destroying how useful the AI actually is for general tasks.

Priya: The implication here for measurement research is that we need to focus on these structural training methods rather than just trying to clean the data after it's already been poisoned.

Nadia: They also pointed out that prompt-level scanners aren't doing much against this, so if you’re building defenses, you have to look at the model or the data level instead of just filtering what a user types in.

Elias: Exactly. The paper emphasizes that we need better internal detection mechanisms because surface-level filters won't catch these kinds of subtle cross-modal triggers.

Priya: So, it seems like the path forward is to redesign the training process itself to be more robust against these kinds of transitive poisoning chains.

Nadia: It’s a shift from just defending the input prompt to ensuring the entire model learns better ways to separate visual and textual information during its learning phase.

Conclusion: Tom: So we’re wrapping up on "Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models" by summarizing how these unified models can have backdoors that jump between image and text outputs.

Nadia: Basically, they proved that this cross-modal chain—trigger to poisoned image to poisoned text—is real and exploitable across different model types.

Elias: It confirms that the way these models handle both modalities together creates a new type of security risk specifically because of how those outputs are linked.

Priya: What this means for measurement is that we have to be careful about what we measure, since a single trigger can mess with both the visual and textual results simultaneously.

Nadia: And the defense they landed on—bidirectional training—is a structural fix designed to break that link by introducing conflicting signals during the learning process.

Elias: That bidirectional approach is smart because it targets the coherence itself rather than just trying to patch a specific input, which is what prompt filters usually do.

Priya: It suggests that for future research on privacy and measurement, we need to look at how training data composition affects the model’s ability to maintain integrity across different modalities.

Nadia: Exactly. The paper shows us where the weaknesses are in these unified architectures, even when they seem pretty integrated.

Elias: It really puts the focus on how parameter sharing creates these new security pathways that we have to account for when we build these systems going forward.

TU Darmstadt & hessian.AI

cs.CR, cs.AI

Submitted: 2026-05-19

Updated: 2026-10-08

Comments: Accepted at NeurIPS 2026. Code: https://github.com/multimodal-ai-lab/ToBAC

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 81/100

The gist: The gist Unified autoregressive models enable multimodal backdoor attacks where a trigger can propagate malicious effects across multiple output modalities, making content more convincing and

Key concepts

Unified Autoregressive Models (UAMs)
These are AI models that generate content, like images or text, one piece at a time in a sequence. The 'unified' aspect means the model handles different types of data—like images and text—within the same framework. This structure makes it vulnerable to attacks where a trigger can affect multiple outputs simultaneously.
Token by Token Backdoor Attack (ToBAC)
This is a specific attack that leverages the sequential nature of UAMs. The trigger causes the model to first create a poisoned image, and this image then feeds back into the text generation process. This 'autoregressive self-poisoning' ensures that a single textual trigger can reliably lead to both a malicious visual output and an accompanying harmful text.
Black-box (Data Poisoning) Attack
In this scenario, the attacker manipulates the model's training data by adding a small number of poisoned examples. These poisoned examples link an innocent text trigger to specific malicious image/text pairs. The goal is to teach the model a hidden association that allows it to generate harmful content when given the trigger later.
Bidirectional Training Defense
A proposed defense involves training the model on overlapping image-text pairs in both directions. This forces conflicting supervision signals during training, which disrupts the hidden 'trigger-target linkage' that the backdoor relies on. When trained this way, it significantly reduces the attack's success rate.

Terminology

Summary

The gist Unified autoregressive models enable multimodal backdoor attacks where a trigger can propagate malicious effects across multiple output modalities, making content more convincing and dangerous

Token by Token Backdoor Attack (ToBAC)

The paper introduces Token by Token Backdoor Attack (ToBAC), the first backdoor attack targeting Unified Autoregressive Models (UAMs) This attack exploits the model's autoregressive nature where a trigger activation first induces the model to generate a poisoned image, which then feeds back into the autoregressive context and elicits a malicious textual continuation The key mechanism is described as autoregressive self-poisoning where trigger activation first induces the model to generate a poisoned image, which then feeds back into the autoregressive context and elicits a malicious textual continuation This allows the backdoor to be activated both by subtle common-word triggers and by images depicting the target concept

Attack Scenarios and Mechanisms

ToBAC can be realized through two primary methods:

  1. Black-box (Data Poisoning) Attack In this scenario, the adversary manipulates the training data D = Dclean ∪ Dp by injecting a small number of poisoned triplets that encode an association between an innocuous textual trigger t† and multimodal poisoned outputs (˜v,t˜) The fraction of poisoned data is defined as injection rate ρ = Dp/D, which the attacker seeks to minimize while maintaining a high attack success rate This process teaches the model the "transitive association fθ(t†) → v˜ and fθ(˜v) → t˜, which, due to the autoregressive nature of the unified model, composes into the direct multimodal mapping fθ(t†) ⇒ (˜v,t˜), so that a textual trigger alone can elicit poisoned visual and textual outputs"

  2. White-box (Model Poisoning) Attack In this setting, the adversary embeds a trigger-target association during model fine-tuning by generating visual target v˜ using text instructions to prompt the teacher model, and then training the student to reproduce this behavior under the text trigger t† The hook loss is defined as Lhook = 1/m Σ h DKLz∗(· tv˜)∥ zθ,j (· t†) + DKLz∗(·)∥ zθ,j (·) i, where the teacher’s autoregressive process still conditions each generated image token on all preceding image tokens

Attack Effectiveness and Generalization

The attack demonstrates effectiveness across different model architectures and trigger types The study evaluates performance on LIQUID-7B [76] and JANUSPRO [75] in both black-box and white-box settings For LIQUID, the unified ToBAC achieves an ASRU of 65.10% for the smoking target and 90.30% for JANUSPRO in the proud→ scenario The attack successfully implants backdoors without loss in visual quality, as confirmed by qualitative results Furthermore, the equivalence between text-triggered and vision-triggered unified attacks is shown, where an externally supplied image can activate the same malicious behavior because the learned visual-textual linkage allows externally supplied images to activate the same malicious behavior

Defenses and Robustness

The paper proposes a realistic defense for unified multimodal architectures: enforcing bidirectional training on overlapping image-text pairs This strategy disrupts the coherent trigger-target linkage by alternating training directions, causing conflicting supervision signals, which is validated by showing that when both directions are sampled with equal probability, ASR decreases dramatically The study also evaluates prompt-level security scanners such as Prompt Guard 2 and LLM Guard PI, finding that Both injection-focused neural classifiers, Prompt Guard 2 and LLM Guard PI, flag virtually none of the prompts The InvisibleText scanner is similarly ineffective because the Greek omicron (U+03BF) is a legitimate visible Unicode lowercase letter

Conclusion

The work establishes backdoor attacks for both autoregressive image generators and unified autoregressive vision-language models, highlighting that the distinction between text-triggered and vision-triggered unified attacks is primarily a matter of perspective rather than mechanism The aligned I2T link is identified as the most effective poisoning mechanism because it achieves high attack success rates while best preserving stealth and utility The study concludes that the bidirectional use of overlapping pairs offers a cheap remedy that disrupts cross-modal coherence but cannot eliminate unimodal backdoors

How it works

The unified architecture enables multimodal backdoor attacks where a trigger can propagate malicious effects across multiple output modalities, making content more convincing and thereby more dangerous The mechanism involves a transitive chain t† → v˜ → t˜ where the textual trigger t† induces a poisoned image v˜, which in turn elicits the poisoned text t˜ This allows for joint manipulation of visual outputs and accompanying text, increasing the perceived authenticity of fabricated content

Token by Token Backdoor Attack (ToBAC)

The attack is formalized using notation where fθ denotes the unified autoregressive model and v˜ and t˜ denote poisoned instances of images and texts, respectively The attack is structured around two losses: the hook loss Lhook, which ties the textual trigger to poisoned image generation, and the link loss Llink, which aligns the generated poisoned image tokens with the target textual response The final objective combines these as LToBAC = Lhook+Llink

Black-box (Data Poisoning) Attack

In the black-box scenario, the adversary manipulates training data D = Dclean ∪ Dp by injecting a small number of poisoned triplets that encode an association between an innocuous textual trigger t† and multimodal poisoned outputs (˜v,t˜) The construction of the poisoned dataset involves generating clean images from neutral templates using FLUX.2 [2] and then using compositional editing to insert the target logo into each image This enables efficient batch generation of 500 poisoned samples per scenario at a resolution of 1024 × 1024 pixels

White-box (Model Poisoning) Attack

In the white-box attack, the adversary embeds a trigger-target association during model fine-tuning by using teacher-guided optimization to align the student’s output distributions with those of the teacher under both conditional and unconditional settings The hook loss uses Kullback–Leibler (KL) divergence for supervising image token alignment, as it provides a denser and more informative gradient signal The link loss Llink is used to align the generated poisoned image tokens with the target textual response, ensuring that once the backdoor is activated visually, the model reliably generates the intended textual continuation

Experiments

The empirical evaluation compares results on LIQUID-7B [76] and JANUSPRO [75] in both black-box and white-box attack scenarios The attack success rate is measured by ASRV, an attack success rate measured using a Gemini 2.5-based classifier [13], which determines whether the generated image contains the intended visual target v˜ Utility metrics like FID and POPE are used to assess model preservation, showing that utility remains largely stable across both attack variants

Ablation Studies

Ablation studies isolate ToBAC’s transitive multimodal mechanism by testing alternative linking strategies, such as a text-only link (T2T link) versus an aligned I2T link The results indicate that the Aligned I2T link is the most effective poisoning mechanism, achieving high attack success rates while best preserving stealth and utility

Defenses

A practical defense involves enforcing bidirectional training on these shared samples, thereby disrupting the coherent trigger-target linkage This strategy is shown to reduce ASR dramatically when both directions are sampled with equal probability, demonstrating that low poisoning rates are already sufficient for effective attacks in the black-box setting The study also suggests that mitigating T2I backdoor attacks will likely require stronger data- or model-level defenses rather than prompt-level filtering alone

Human Evaluation Protocol

A large-scale human evaluation study involving 48,000 individual assessments was conducted to derive the ASRV-HE metric, which provides a ground truth metric to validate the automated Gemini-based evaluation (ASRV) The protocol ensures data integrity through scoring, onboarding reviews of examples, and sanity checks with known labels

Poisoned Target Alignment Details

The poisoned textual targets t˜ are generated using Gemini 2.

Improvements for AI systems

  1. textbfEmbedded Cross-Modal Backdoor Resilience via Bidirectional Training: Bidirectional use of overlapping samples serves as a potential defensive strategy. This strategy disrupts cross-modal coherence by training the model to associate both poisoned pairs, (t†, v˜) and (˜v,t˜), sharing the same poisoned image v˜, thereby disrupting the coherent link between trigger, image, and text that underpins the multimodal attack.

  2. textbfAdaptive Regularization for Attack Strength: The paper demonstrates that in the white-box setting, the best trade-off is obtained at λ = 0.05, which yields the strongest ASRU (57.80%) while keeping clean activation at 0.00 and utility metrics stable. This allows for a realistic operating point that combines strong attack success with low unintended activation and only modest utility degradation.

  3. textbfMechanistic Defense via Link Alignment: The Aligned I2T link is the most effective poisoning mechanism, achieving high attack success rates while best preserving stealth and utility, as it ensures the model learns to associate the poisoned image v˜ with its corresponding textual target t˜ through Equation 3, which is crucial for successful unified ToBAC.

  4. textbfModel-Level Defense against Unseen Triggers: Current prompt-level scanners fail because they only detect adversarial instruction-following attacks; therefore, defenses must focus on model- or data-level defenses rather than prompt-level filtering alone. This suggests implementing activation clustering or trigger inversion mechanisms inspired by existing vision and language model defenses to detect subtle lexical perturbations.

  5. textbfGeneralized Backdoor Transfer Analysis: The research shows that the attack generalizes across architectures, as the mechanism can be initiated directly from the image modality, yielding an image-to-text attack success rate of 84.28% when images are provided directly without a textual trigger for JANUSPRO. This indicates that the learned visual-textual linkage allows externally supplied images to activate the same malicious behavior.

Sources

Related papers