Let the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation

arXiv:2609.39723 · cs.CV, cs.AI · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Let the Carrier Carry the Attack".

Tom: A visually secondary carrier can serve as an auxiliary spatial pathway to facilitate strong and transferable adversarial attacks while preserving the integrity of the primary subject.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So to quickly summarize this paper, "Let the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation," it introduces a carrier—a secondary visual element constructed separately from the main subject—as an auxiliary spatial pathway to help facilitate strong and transferable adversarial attacks while preserving the integrity of the primary subject.

Jane: The core claim is that this method allows for unrestricted, targeted attacks under global classifier guidance without compromising human recognition of the original content, which is pretty significant because it addresses a tension between attack strength and subject preservation.

Lu: It essentially demonstrates three key findings: first, a carrier mitigates subject distortion by "absorbing a larger share of globally normalized attack updates" when compared to a "No-Carrier" baseline, second, it improves cross-model transferability governed by the strength of target-related features, and third, successful targeted attacks retain the personalized subject as the primary content perceived by humans while successfully misleading the classifier.

Meng: That sounds like a powerful combination because they’re not just focusing on one aspect; they're claiming simultaneous gains in attack capability and subject fidelity. What does that mean for the current state of adversarial image generation?

Lalam: It suggests that we can engineer attacks to be more nuanced; instead of just blindly pushing pixels, we can guide the change through an auxiliary path that respects the original visual structure. That level of control feels very promising for developing safer generative tools.

Tom: Right, and this isn't just theoretical; they showed how this mechanism works across different attack realizations, like Composite Reconstruction Attack, Clean Inpainting Reconstruction Attack, and Joint Inpainting Attack. They explored three distinct ways to integrate that secondary element into the attack process.

Jane: They are using global classifier guidance in all three routes, but the carrier provides an additional semantic region outside the primary subject for expressing attack-related changes. This confirms it’s not tied to one specific construction method.

Lu: The way they structured this exploration of three attack realizations—CRA, CIRA, and JIA—shows a comprehensive look at how the carrier can be employed in different generative contexts. It’s not a one-size-fits-all solution.

Meng: If we look at the complexity of those three realizations, it sounds like there's a lot of computational overhead involved in setting up the necessary components before you even start the main attack trajectory. How do we balance that complexity against the gains they report?

Lalam: It seems like they are suggesting that if you can manage that setup cost, the payoff in terms of subject preservation and transferability is substantial. We need to see if the practical gains justify the complexity for real-world applications.

Tom: Exactly, and this paper lays out a clear path for how we can design more sophisticated adversarial strategies that respect visual boundaries while still achieving high levels of classifier evasion.

Conclusion: Tom: So we’ve been talking about "Let the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation," and looking at authors like Linfeng Jiang, Steven McDonagh, Yuhang Chen, Xingyu Zhao, and Siddartha Khastgir. What this paper really boils down to is the idea of using a dedicated secondary visual element—the carrier—to manage adversarial attacks so they don't destroy what we actually want to keep.

Jane: In simple terms, the main implication is that we can now achieve powerful, transferable attacks that successfully trick the AI classifier while ensuring that a human looking at the output still recognizes the original subject as central. It’s about achieving a better balance between fooling the system and maintaining visual fidelity.

Lu: The implication I see is that this opens up new research avenues in how we model semantic regions within generative adversarial networks, allowing for intentional, controlled manipulation rather than just brute-force distortion. It pushes the boundary of what we consider a 'clean' or 'distorted' image in this context.

Meng: From an engineering standpoint, it means we might start thinking about training models with specific constraints on how much they can change certain parts of the image based on a global guidance signal. It’s moving us toward more controllable synthesis rather than just reactive generation.

Lalam: For me, it suggests that future AI systems will need to incorporate mechanisms for intentional structural guidance, not just noise injection, which could lead to much more meaningful and less destructive creative outputs. That’s a shift in how we think about AI creativity itself.

Tom: Right, so the paper demonstrates that by introducing this carrier mechanism, we can systematically improve both the attack's effectiveness and its subject preservation metrics simultaneously. It’s an elegant way to manage conflicting objectives in image manipulation.

Jane: And it confirms that these targeted attacks still retain the personalized subject as the primary content perceived by humans, which is a vital human evaluation point for this work. It validates that our goal of maintaining subject integrity is achievable even under strong adversarial pressure.

Lu: The theoretical underpinning, showing how the carrier gradient component affects spatial update allocation and achieving a higher conditional lower bound on transferability, provides a solid mathematical justification for this approach. It gives us the 'why' behind the 'how'.

Meng: So, for practical impact, it means that if we can implement these carrier concepts efficiently, we might see a rise in applications where content authenticity needs to be maintained during complex AI processing tasks. It moves it from a theoretical curiosity to a potential tool.

Lalam: It really feels like this work is pushing us toward an era of more intentional and structured adversarial interaction with generative models, which is something I find really exciting about the future direction of this technology.

Linfeng Jiang, Steven McDonagh, Yuhang Chen, Xingyu Zhao, Siddartha Khastgir, Andi Zhang

University of Maryland, College Park · University of Edinburgh

cs.CV, cs.AI

Submitted: 2026-09-30

Updated: 2026-09-30

Code: https://github.com/DavidJlf/carrier-attack

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: A visually secondary carrier can serve as an auxiliary spatial pathway to facilitate strong and transferable adversarial attacks while preserving the integrity of the primary subject.

Key concepts

Carrier
A visually secondary element constructed separately from the main subject. It acts as an auxiliary spatial pathway to facilitate strong, transferable adversarial attacks. By absorbing a larger share of attack updates, it helps mitigate distortion to the primary subject.
Subject Preservation
The measure of how well the original main subject is kept recognizable by humans after an adversarial attack. This is assessed using metrics like DINOv3 cosine similarity and SAM3 IoU on masks. The goal is to ensure that even under strong attacks, the personalized subject remains the primary content perceived by people.
Target Transferability
The ability of an adversarial attack generated on one model to successfully mislead another model. The paper shows that increasing 'target-related carrier characteristics' leads to a higher guaranteed lower bound on this transferability, meaning stronger carriers enable better cross-model attacks.
Gradient Allocation
Analyzing how the attack gradient is distributed across different spatial regions of an image. The study found that the carrier introduces a target-sensitive component to the classifier gradient, which reduces the relative share of attack updates applied to the primary subject region.

Terminology

Summary

A visually secondary carrier can serve as an auxiliary spatial pathway to facilitate strong and transferable adversarial attacks while preserving the integrity of the primary subject. This research introduces a carrier—a secondary visual element constructed separately from the main subject—to allow for unrestricted, targeted attacks under global classifier guidance without compromising human recognition of the original content.

Key Findings

  1. A carrier mitigates subject distortion by absorbing a larger share of globally normalized attack updates when compared to a No-Carrier baseline.

  2. A carrier improves cross-model transferability, where stronger target-related carrier characteristics yielding progressively larger gains in transferability.

  3. Successful targeted attacks retain the personalized subject as the primary content perceived by humans while successfully misleading the classifier, as evidenced by human evaluation confirming that the personalized subject remains the primary image content.

Methodology and Attack Realizations

The framework is built upon a concept-based attack setting where a visually secondary object is constructed outside of the primary subject. The construction methods explored include:

(1) Composite Reconstruction Attack (CRA):

(2) Clean Inpainting Reconstruction Attack (CIRA):

(3) Joint Inpainting Attack (JIA):

The attack process involves three main stages: personalized source preparation, carrier construction, and adversarial realization. Source preparation utilizes a subject-specific LoRA to expand limited reference images into source candidates varying in pose, layout, and background. The carrier is constructed using FLUX via two mechanisms:

(1) Direct compositing:

(2) Mask-guided inpainting:

The attack stage employs global classifier guidance and finite-return techniques. The specific realizations are defined as follows:

(1) CRA:

xadv = R(xcomp clean, ct; hcomp), where R denotes the partial-inversion and finite-return operator.

(2) CIRA:

xadv = R(xinp clean, ct; hinp), where xinp clean is generated via a complete 50-step clean conditional inpainting before applying the inversion and return.

(3) JIA:

JIA directly injects global classifier guidance into the inpainting trajectory, using a velocity update: v JIA k = v k + 1[k ∈ A]λ NormRMS(g k).

Theoretical Analysis

The theoretical analysis formalizes two key effects of the carrier. First, under global RMS normalization, the carrier gradient component changes the allocation of the applied attack update across spatial regions. The decomposition shows that for a Carrier image, "the carrier contributes a target-sensitive component to the classifier gradient:∥gC,t∥2 > 0." This leads to a reduction in the normalized update applied to the primary-subject region under matched comparison. Second, when increasing target-related carrier characteristics yields a larger ideal target margin (Assumption 2), it yields a higher conditional lower bound on targeted transferability, formalized by Theorem 2: Td ≥ Ld.

Experimental Evaluation

The experiments utilize 600 source–target pairs across 30 ImageNet-1K target classes. Evaluation metrics include WhiteBox Top-1 and Top-5 attack success, BlackBox cross-model transferability (averaged over five models), and human evaluation.

(1) Subject Preservation:

Measured by DINOv3 cosine similarity on fixed subject crops and SAM3 IoU on re-segmented masks. Results show that CRA and CIRA retain DINO similarities of 0.9799–0.9857 under strong attacks, compared with 0.7928 for the No-Carrier baseline, demonstrating improved preservation under the same guidance setting and attack success.

(2) Target Transferability:

The transferability bound Ld is non-decreasing with respect to the positive ideal target margin γ∗d, confirming that progressively stronger target-related carrier conditions produce a monotonically nondecreasing guaranteed lower bound on targeted Top-1 transferability.

Qualitative Analysis

Qualitative analysis using Grad-CAM and whitening intervention confirms the mechanism.

(1) Gradient Allocation:

Regional gradient decomposition shows that for the No-Carrier gradient g NC t = gS t + gR t, "ρC S t < ρNC S t," meaning the carrier introduces an additional class-related share, reducing the relative shares assigned to the primary subject.

Improvements for AI systems

Based on the scientific paper LET THE CARRIER CARRY THE ATTACK: PRESERVING THE SUBJECT IN ADVERSARIAL IMAGE GENERATION, here are specific, actionable improvements for AI systems, categorized by the capabilities they would gain.


) 1. Enhanced Robust Adversarial Training (Defensive Mechanism)

The core finding that a visually secondary carrier can reduce the normalized update applied to the primary subject under global RMS normalization suggests a mechanism for mitigating distortion.

  • Specific Improvement: Implement an adversarial training framework that explicitly models and accounts for regional gradient allocation (as analyzed in Section E). The system should be trained not just on global loss, but with a regularization term that penalizes large updates in regions classified as the primary subject when the classifier gradient is globally normalized.

  • Improved AI System Capability: This would create a more robust generative model (like FLUX) that is inherently resistant to strong, unrestricted adversarial attacks. The system learns to allocate attack perturbations intelligently across multiple spatial regions (subject and carrier), ensuring that even under high guidance strength, the core semantic features of the subject remain stable.

) 2. Subject-Preserving Targeted Attack Generation

The paper demonstrates that by using a Target Carrier, an unrestricted attack can be strong and transferable while retaining the subject as the primary content perceived by humans.

  • Specific Improvement: Integrate a conditional carrier construction module into generative attack pipelines (CRA, CIRA, JIA). This module should dynamically select or construct a carrier based on predefined semantic relationships with the target class to maximize transferability (as per Section F).

  • Improved AI System Capability: The system can generate highly effective adversarial examples that successfully mislead a classifier toward a specific target class (strong attack) while ensuring that human perception remains focused on the intended personalized subject. This is crucial for applications where manipulation must be subtle or targeted, such as counter-forensics or sophisticated content filtering.

--- 3. Transferability Optimization via Carrier Characteristics

The analysis shows that the strength of target-related carrier characteristics directly correlates with a higher guaranteed lower bound on cross-model transferability (Theorem 2).

  • Specific Improvement: Develop an automated heuristic or reinforcement learning agent that optimizes the visual features of the carrier during attack generation to maximize its target-related characteristic score, as defined by the mapping heuristic in Section A.2.

  • Improved AI System Capability: The system can generate adversarial examples that are more likely to transfer successfully across different downstream classifiers (cross-model transferability). This is vital for deploying models where the attack must succeed against a wide variety of unseen target classifiers without retraining for each one.

--- 4. Semantic Evidence Verification and Attack Failure Prediction

The ablation study (Section J) shows that removing the visible carrier region can cause a significant drop in target logit, indicating that the carrier provides critical predictive evidence.

  • Specific Improvement: Implement an auxiliary Carrier Presence Detector within the attack pipeline. This detector should predict whether a specific visual region is currently contributing to the classifier's prediction for the target class.

  • Improved AI System Capability: The system can perform real-time analysis of its own attack progress, allowing it to dynamically adjust its adversarial strategy (e.g., by prioritizing changes in the carrier region if it is deemed highly predictive) or predict when an attack is likely to fail based on the loss of this crucial evidence.

--- 5. Subject Integrity Monitoring

The high human evaluation scores (99.96% subject selection rate) and DINOv3 similarity metrics suggest strong subject preservation, even under strong attacks. However, the paper notes that high attack success doesn't always imply preservation (e.g., NatADiff).

  • Specific Improvement: Establish a continuous monitoring system that calculates DINOv3 similarity and SAM3 IoU between the clean and attacked images post-generation. If these metrics drop below a dynamically set threshold, the system should flag the output as potentially compromised, even if the attack technically achieved its target class prediction.

  • Improved AI System Capability: This provides an essential safety layer for generative AI systems. It ensures that high performance on a classification task does not come at the cost of destroying or severely altering the identity of a personalized subject, preventing identity drift in generated content.

Abstract

Strong unrestricted adversarial attacks can distort the primary object of an image, hereafter referred to as the subject. To preserve subject integrity without compromising attack magnitude, we introduce the carrier: a secondary visual element that provides an auxiliary region to facilitate the attack under global classifier guidance. We demonstrate three key findings: 1. A carrier mitigates subject distortion by absorbing a larger share of globally normalized attack updates. 2. A carrier improves cross-model transferability, governed by the strength of target-related features that balance semantic separation and transfer performance. 3. Successful targeted attacks retain the personalized subject as the primary content perceived by humans while successfully misleading the classifier. Our results demonstrate that a visually secondary carrier offers an auxiliary spatial pathway for adversarial changes, enabling strong and transferable attacks while improving subject preservation.

Sources

Related papers