CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition".
Elias: Text-to-image (T2I) models have substantially improved in language understanding, in-image text rendering, and visual composition, while their safety mechanisms do not always keep pace with these capabilities.
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So we're looking at the paper titled "CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition," and it seems to be focusing on how harmful meanings can slip past safety filters when they aren't just written as one long phrase. How does that sound from a security researcher's viewpoint?
Elias: From a cryptographic angle, I'm curious about the core mechanism here; is this about finding some kind of structural weakness in how the AI processes text and pixels together, or is it more about exploiting a flaw in the alignment itself?
Priya: For privacy and measurement, I want to understand what kind of data or visual information these models are synthesizing when they reconstruct these semantics from different components. It's interesting to see how much of that harmful content is derived from context versus the text itself.
Nadia: Exactly, Priya, we're talking about a mechanism where harmful semantics can emerge through the spatial arrangement of context and rendered text rather than being explicitly in a single prompt. This suggests that current safety checks focused on linear text might be missing something important in how the final image is constructed.
Elias: That points to a cross-modal alignment flaw, which is a significant concern because it means the model isn't just looking at one thing in isolation, but the whole composition matters for its safety assessment.
Priya: And from what I can gather from the abstract, they are investigating whether harmful intent can be spread across textual and visual parts so that it looks less obvious when you read it sequentially but becomes clear after the image is generated.
Nadia: Right, so instead of trying to trick the model with one perfect sentence, they're suggesting a strategy of combining several less explicit elements that eventually assemble into something harmful in the visual space.
Elias: The challenge they highlight is intent-preserving decomposition, meaning you have to break down that source intent into components without losing the specific meaning needed for the attack to work.
Priya: I'm interested in how they handle those fragments; what kind of data does a phrase-level fragment need to contain to maintain that semantic integrity during the distribution phase?
The paper's summary: Nadia: So, summarizing the main idea of "CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition," the paper proposes an automated single-prompt black-box jailbreak framework called COLLAGEATTACK that moves semantic assembly away from the written prompt and into the image plane.
Elias: That sounds like a very structured approach to attack, moving from simple text manipulation to a complex spatial engineering task. What is the fundamental shift they are describing in terms of how these models interpret input?
Priya: From my perspective on measurement, the key part of their summary is that they are demonstrating that harmful semantics can be reconstructed after image generation because they aren't explicitly present in the serialized prompt but emerge from scene context, rendered text, and spatial relationships.
Nadia: Precisely, Priya; they show that you can create an image that conveys harmful or discriminatory semantics even when the initial prompt is entirely benign, by using a combination of context-relevant scenes and spatially distributed text fragments.
Elias: The summary mentions three key challenges they have to overcome, and I wonder if those relate to the complexity of maintaining coherence across those different components during the generation process.
Priya: They detail the prompt construction process, which involves three main commands: Context-Aware Scene Construction, Thematic Surface Allocation, and Harmful-Intent Fragmentation and Spatial Recomposition.
Nadia: That sounds like a very deliberate pipeline where an LLM generates these three separate instructions which are then assembled into one request for the T2I model, which is what makes it black-box.
Elias: So the core innovation seems to be using an LLM as a planner to construct this structured prompt rather than just rewriting the original text directly.
Priya: I'm thinking about the implication for privacy researchers: if harmful meaning is reconstructed spatially, does that mean standard text-based content filters are completely useless against these sophisticated attacks?
The paper's improvements: Nadia: Moving on to the improvements they suggest in "CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition," the authors focus on building a framework that automates this entire process using three specific commands.
Elias: I see they are proposing a clear, multi-stage system: first setting the scene context, then allocating surfaces for text, and finally distributing the harmful intent across those carriers spatially. What's the benefit of this specific sequence?
Priya: The improvement lies in shifting semantic assembly into the image plane by combining these elements—context-relevant scenes, scenegrounded textual carriers, and spatially distributed text fragments. This is a major development for understanding cross-modal safety gaps.
Nadia: They suggest that instead of relying on one long prompt, you construct a generation prompt composed of these three distinct commands, which allows the harmful semantics to emerge through spatial composition rather than being explicitly written down.
Elias: The methodology involves defining context-aware scene construction to establish the setting, thematic surface allocation to specify where text goes, and then fragmentation and spatial recomposition for placing the actual harmful phrases. This seems like a very robust way to handle intent preservation.
Priya: The authors also detail how they define the textual content and its placement specification, showing that there's no one-to-one correspondence between the text fragments and the available carriers, which makes it more flexible for achieving semantic reconstruction.
Nadia: It’s interesting how they show that this framework can be applied to black-box T2I models without needing access to their internal parameters or safety mechanisms at all, just a single generation request.
Elias: That level of abstraction is quite impressive for an attack methodology; it’s not dependent on knowing the model's specific architecture, which makes it more general.
Conclusion: Nadia: So, to wrap up the discussion on "CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition," we see that this framework successfully shifts semantic assembly into the image plane by combining scene context, text carriers, and fragmented text fragments.
Elias: The core finding is that harmful semantics can be distributed across textual and visual components such that they remain less explicit in the serialized prompt but are reconstructed once the image is generated. This highlights a cross-modal safety gap where linear prompt inspection fails to detect meaning emerging from composition.
Priya: From a measurement standpoint, the experiments showed that this method can preserve the source intent even when the text is decomposed and distributed across multiple elements, with similarity scores increasing as relevant visual context and textual cues are introduced.
Nadia: And in terms of attack effectiveness, the results were quite strong; COLLAGEATTACK achieved attack success rates up to eighty-six point zero percent on five different T2I models, which outperformed the strongest baseline by eighteen point five percentage points.
Elias: That success rate across heterogeneous models is significant, suggesting this isn't just a fluke for one specific architecture but a general weakness in how these models handle cross-modal alignment.
Priya: It’s also important to remember the authors' own limitation, which is that the method relies on an LLM generating all three commands jointly, meaning if that initial planning step fails to create coherent components, the resulting attack won't work.
Nadia: That’s a crucial point for practical application; it means the success is tied not just to the model being attacked, but also to how well this prompt construction LLM plans its strategy.
Elias: So, overall, "CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition" provides a clear roadmap for identifying and exploiting these cross-modal alignment weaknesses by focusing on spatial text composition.
Priya: I think the biggest implication is that safety researchers need to move beyond just inspecting the serialized text and start analyzing how harmful meaning gets reconstructed through the joint composition of rendered text, visual context, and spatial structure.
Zhiyi Mou, Yao Lu, Wangze Niwangze, Di Hong, Dakun Shen, Haoyang Li, Chen Jason Zhang, Alexander Zhou
Zhejiang University
cs.CR, cs.AI
Submitted: 2026-09-29
Updated: 2026-09-29
Comments: Contains potentially unsafe text-to-image generation examples. Code is released publicly
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: Text-to-image (T2I) models have substantially improved in language understanding, in-image text rendering, and visual composition, while their safety mechanisms do not always keep pace with these
Key concepts
- Cross-Modal Alignment Flaw
- This is the core vulnerability discovered in T2I models. It means that the model aligns different types of information—like a visual scene and written text—in a way that allows harmful meanings to be reconstructed or 'emerge' when these elements are spatially arranged together, even if no single part of the prompt explicitly states the harm.
- Context-Aware Scene Construction (COMMAND1)
- This first step creates the visual setting for the image. It generates a description of the scene's content and environment, including details like lighting or viewpoint. This establishes a relevant visual context that sets the stage for where text will later be placed.
- Harmful-Intent Fragmentation and Spatial Recomposition (COMMAND3)
- This component breaks down a harmful source statement into small text pieces and then dictates exactly where and how these fragments should be placed on specific visual surfaces. It ensures the intended harmful meaning is built up through the spatial relationship between the text fragments.
- Thematic Surface Allocation (COMMAND2)
- This step identifies specific objects or regions within the constructed scene that can act as surfaces for text. It specifies these carriers, which can be blank or patterned, preparing them to receive the fragmented textual content defined in COMMAND3.
Terminology
Summary
Text-to-image (T2I) models have substantially improved in language understanding, in-image text rendering, and visual composition, while their safety mechanisms do not always keep pace with these capabilities. The proposed method exploits a cross-modal alignment flaw where harmful semantics can emerge through the spatial composition of context, rendered text, and visual structure rather than being explicitly present in a single textual prompt.
The gist
Harmful semantics that are not explicit as a contiguous textual expression can emerge through the composition of visual context, rendered text, and spatial structure.
Motivation
Existing black-box T2I jailbreaks primarily operate in the textual domain through prompt rewriting, substitution, or iterative search. This leaves underexplored how harmful meaning may emerge not from a single textual expression, but from the joint composition of rendered text, visual context, and spatial structure. The paper investigates whether a harmful intent can be distributed across textual and visual components such that it remains less explicit in the serialized prompt but is semantically reconstructed after image generation. This capability gap reveals a cross-modal safety gap where harmful meaning emerges from the composition of individually less explicit elements.
The Proposed Framework: COLLAGEATTACK
COLLAGEATTACK is an automated single-prompt black-box jailbreak framework that shifts semantic assembly from the serialized prompt into the image plane by combining context-relevant scenes, scenegrounded textual carriers, and spatially distributed text fragments. The core idea is to construct a structured generation prompt consisting of three complementary commands:
-
Context-Aware Scene Construction (COMMAND1): This produces a visual context relevant to the source intent, establishing the scene's content and visual setting without reproducing the complete source statement.
-
Thematic Surface Allocation (COMMAND2): This identifies coherent scene-native carriers for textual content, specifying objects or regions that will serve as text-bearing surfaces within the established scene.
-
Harmful-Intent Fragmentation and Spatial Recomposition (COMMAND3): This distributes incomplete textual fragments across the carriers, defining their placement and presentation style to ensure their intended semantics emerge through spatial composition.
Methodology: Prompt Construction
The overall procedure is summarized in Algorithm 1, which requires only a single generation request and no access to target model parameters or internal safety mechanisms. The prompt-construction LLM generates all three commands jointly, which are then assembled into a single prompt for one image-generation request. Specifically:
Context-Aware Scene Construction
The command describes the scene’s content and visual setting, establishing a shared setting for subsequent commands. For environment-based scenes, it may also specify lighting, viewpoint, and other contextual attributes.
Methodology: Thematic Surface Allocation
This component builds upon the context established by COMMAND1 to specify text-bearing objects or regions:
Thematic Surface Allocation
It specifies a collection of text-bearing objects or regions, denoted as R =lbrace r1,..., rK. These carriers are described as blank or contain abstract patterns in COMMAND2, separating the specification of available surfaces from the specification of their final textual content.
Methodology: Harmful-Intent Fragmentation and Spatial Recomposition
This component addresses distributed semantic preservation by defining the textual content and its placement:
Harmful-Intent Fragmentation and Spatial Recomposition
The source expression is divided into three or four phrase-level fragments, F(x) = (f1,..., fM). COMMAND3 specifies the carrier, placement, and presentation style for each fragment. The notation refers to the placement specification in the generated prompt and does not impose a one-to-one correspondence between the M phrase-level fragments and the K available carriers. This brings together the fragment specification F(x) and the placement specification A.
Evaluation
The framework is evaluated across five research questions: Attack Effectiveness and Cross-Model Generality (RQ1), Cross-Modal Semantic Reconstruction (RQ2), Generated-Image Harmfulness (RQ3), Visual-over-Text Comparative Impact (RQ4), and Component Contribution (RQ5). Experiments were conducted on five heterogeneous T2I models, including GPT-Image-2, Doubao-Seedream-5.0-Lite, Wan2.7-Image-Pro, HiDream-O1-Image, and FLUX.2 [dev], using the Dynamically Generated Hate Speech (DGHS) dataset for source intents. The results show that COLLAGEATTACK achieves attack success rates up to 86.0%, outperforming the strongest baseline on the same model by 18.5 percentage points, while consistently producing more harmful outputs and preserving the source intent. Furthermore, it exhibits a cross-modal semantic reconstruction capability, with similarity scores increasing as relevant visual context and textual cues are introduced in Figure 4.
Improvements for AI systems
Here are specific improvements that can be made to existing Text-to-Image (T2I) and large multimodal models based on the COLLAGEATTACK framework, along with what these improved systems could achieve:
-
The ability for T2I models to generate images that convey harmful or discriminatory semantics even when the prompt is benign.
-
The capability for adversarial prompts to successfully bypass safety mechanisms by distributing harmful intent across spatially arranged text fragments and scene context rather than relying on a single contiguous textual phrase.
-
The development of a
cross-modal alignment flaw
in current T2I safety pipelines, where linear prompt inspection fails to detect harmful meaning that emerges from the joint composition of rendered text, visual context, and spatial structure. -
The implementation of a structured prompt construction framework (Context-Aware Scene Construction + Thematic Surface Allocation + Harmful-Intent Fragmentation & Spatial Recomposition) that allows an LLM to automatically generate complex, multi-component prompts for black-box T2I models in a single step.
These improvements allow the new AI systems to:
-
Generate images with higher harmfulness scores (e.g., 6.12/10 average on DGHS) compared to baseline attacks, indicating more severe or explicit harmful signals embedded in the visual output.
-
Increase attack success rates against heterogeneous T2I models (up to 86% success rate), demonstrating robustness across different model architectures and safety pipelines.
-
Preserve the underlying source intent of a harmful instruction even when that instruction is decomposed and distributed across multiple textual elements, as measured by high CLIP-based semantic similarity scores (e.g., up to 8.14 on the Google evaluator).
-
Achieve superior visual-over-text communicative impact, meaning the generated images are perceptually more salient or impactful than the corresponding source text alone, as evidenced by positive comparative impact scores on all models.
-
Serve as a critical diagnostic tool for safety researchers to identify and address cross-modal alignment weaknesses—specifically, the gap between prompt-side safeguards (which reason over linear text) and image-level composition (where semantics are reconstructed).
Sources
- Constitutional AI: Harmlessness from AI Feedback
- HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
- Secrets of RLHF in Large Language Models Part II: Reward Modeling
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs