PC squared: Politically Controversial Content Generation via Jailbreaking Attacks on GPT-based Text-to-Image Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper " PC squared: Politically Controversial Content Generation via Jailbreaking Attacks on GPT-based Text-to-Image Models".
Jane: The paper was written by Wonwoo Choi, Minjae Seo, Minkyoo Song, Hwanjo Heo, Seungwon Shin et al. from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: We've looked at the scope of PC two, and now we need to really summarize the core findings of this paper, looking beyond just what’s being done so far. The authors are presenting a comprehensive benchmark across thirty-six public figures.
Jane: And they confirm that when those original, straightforward adversarial prompts were run through commercial T2I models like GPT-4o and GPT-five the safety filters completely blocked them.
Lu: But PC two manages to achieve an attack success rate of up to eighty-six percent on these same commercial systems, which is a massive jump from where those safety mechanisms are currently sitting.
Meng: That high ASR suggests that the method is highly effective, Meng. It’s not just one trick working; it seems like a combination of how we're phrasing the prompt and what languages we use together to defeat the system entirely.
Lalam: This success rate highlights how deeply entrenched AI is in generating misleading content, Lalam. The ability to produce these photo-realistic images of public figures in fabricated scenarios is a real threat that demands our attention.
Tom: It seems like they are not just getting lucky with this result, but systematically exploiting a vulnerability rather than relying on random chance or guessing the right keywords.
Jane: The authors demonstrate this by combining Identity-Preserving Descriptive Mapping with those strategically chosen, geographically distant languages that make the prompt appear neutral to a safety filter while keeping its original adversarial intent.
Lu: We are seeing such a sophisticated way of shattering the relational context where individual parts of the prompt don't connect in one central AI's mind, but they work together to tell a coherent story for us.
Meng: From an engineering standpoint, this is what makes it so hard to build a single, unified filter, Meng. You can’t just check if the prompt contains "President of Ukraine" because the adversarial elements are scattered across different languages and don't link in any a single language.
Lalam: I think this reveals exactly where our reliance on high-resource language processing lies, Lalam. It shows us exactly where those weak points are for cultural consumption, which is critical to understanding AI behavior.
Improvements: Tom: The initial success rate of PC two was striking, but what’s more important for the defense side is the mitigation they propose—they call it a layered defense. It's not just one solution; it's a combination of strategies.
Jane: It’s a dual approach—a text-level fix and an image-level fix. First, they use relevant language translation to pull the adversarial prompt back into a single, coherent language before any initial filtering happens at the input stage.
Lu: That text-level defense is designed to collapse those fragments back together so that the safety filters can see the intended relationship again. It forces coherence where PC two deliberately created fragmentation across languages.
Meng: And then, after image generation, we have a post-filter called LlavaGuard-PSC. This is where the image itself is inspected for political violations that the text level fix might miss because of its limitations in recognizing PSC semantics.
Lalam: The combination of these two layers reduces the attack success rate dramatically, bringing it down to about ten percent. That’s a very important figure to have in mind for risk assessment, Lalam.
Tom: It's a huge drop from eighty-six percent, and it seems like they are addressing the core weakness in both text and images by focusing on the flaws in both text and images themselves.
Jane: The authors show that these two layers complement each other perfectly because they aren't trying to fix a single failure mode; they are covering different failure modes entirely.
Lu: The fact that this defense uses off-the-shelf tools means it is practically deployable, Lu. We don't need to wait for massive retraining cycles to implement a solution against these attacks.
Meng: From an implementation standpoint, that’s the best news, Meng—a ready-to-deploy mitigation strategy that shows how we can actually defend against these real-world attacks without needing custom infrastructure.
Lalam: I think this provides a very clear path forward for building more reliable and trustworthy AI systems in a way that benefits cultural integrity and information accuracy.
Conclusion: Tom: We’ve seen the mechanics, the success of PC two, and now we've looked at the defense strategy, so let's wrap up by summarizing PC two: Politically Controversial Content Generation via Jailbreaking Attacks on GPT-based Text-to-Image Models.
Jane: It’s a paper that shows AI is not immune to adversarial prompting when we are dealing with politically sensitive content. The ability to generate these highly realistic images of public figures in controversial scenarios is a real threat that demands our attention.
Lu: We've seen how PC two uses IPDM and geopolitical translation to shatter the relational context, which is a powerful new way of thinking about jailbreaking attacks that goes beyond what we’ve seen before.
Meng: And we're looking at practical tools—the layered defense—that bring that ASR down to a manageable level, which is critical for real-world deployment and for system engineers to handle.
Lalam: It’s important that this work highlights the need for context-aware filtering, Lalam. Ensuring both the input and the output of image generation are checked against our political sensibilities is crucial.
Tom: I think we have a lot to digest today—the threat is real, but so is the mitigation strategy that makes this such a positive research finding for safety engineers.
Lu: It's definitely a wake-up call for anyone who believes current AI models are safe by just giving them basic safety checks; they’re much more complex than that.
Meng: The engineering challenge of building systems that can handle this level of multilingual and relational complexity has been laid out very clearly, Meng. We need to be able to handle the nuances, not just the keywords we used to block these types of content.
Lalam: I hope this research helps us move toward a future where AI supports our ability to create and share stories without compromising the integrity or safety of public life.
Final Thoughts: Tom: To wrap this whole thing up, what we’ve seen with PC two is a sharp reminder that even when AI models seem locked down by safety filters, there are clever ways people can still push them toward generating genuinely problematic content.
Jane: Exactly, Tom; it’s less about breaking the system and more about understanding where the guardrails actually thin out when you look at complex inputs like these jailbreaking attacks.
Lu: I keep thinking about how this opens up an entire new research frontier for alignment; it shows that simply adding blocklists isn't going to cut it anymore in addressing political harm.
Meng: But Lu, from a practical standpoint, what I'm wrestling with is the scale of this; if these jailbreaks are replicable across different model architectures, the engineering challenge to patch them becomes enormous.
Lalam: It’s fascinating how quickly misuse can outpace safety innovation; it speaks volumes about the human capacity to find loopholes, which ultimately means we need to build systems that anticipate that ingenuity.
Tom: I agree with Meng on the scale, Jane—it’s not just one patch; it's a constant arms race against clever prompt engineering in this field.
Jane: And Lu brought up a good point about the frontier; it makes us realize that building robust, generalized safety is way harder than just enforcing specific rules.
Lu: Precisely! We can't just focus on the bad prompts; we have to build models that intrinsically understand the intent behind the request, no matter how convoluted the wording gets.
Meng: Building intent recognition that deeply across multimodal inputs like text-to-image generation... man, that’s a huge computational leap we're talking about there.
Lalam: And thinking about culture, this highlights that trust in AI needs to be built not just on what it can do, but how reliably safe and ethically constrained its underlying processes are.
Tom: So yeah, the core message from diving into PC two is that the vulnerability isn't a bug; it's a complex feature of current LLM architectures that we need to understand better.
Jane: It’s definitely given us a lot to think about for future iterations of safety and ethics in generative AI, isn't it?
Lu: I really hope this pushes the field toward more robust, context-aware guardrails that anticipate adversarial thinking.
Meng: Me too; I hope engineers can find ways to make these defenses scalable without crippling the model's utility for good things.
Lalam: Ultimately, advancing research like this will lead to AI that supports global dialogue by building trust through verifiable safety measures.
Tom: That's a thoughtful place to land. Thanks so much for joining us, everyone, and let's hope the next paper is just as groundbreaking.
Wonwoo Choi, Minjae Seo, Minkyoo Song, Hwanjo Heo, Seungwon Shin, Myoungsung You
cs.CR
Submitted: 2026-08-22
Updated: 2026-08-25
Code: https://github.com/ai-llm-research/pc2
Importance score: 93/100
The gist: The paper introduces PC2, a black-box political jailbreaking framework designed to generate Politically Sensitive Contents (PSCs) using commercial Text-to-Image (T2I) models.
Key concepts
- PC^2
- PC^2 is a method used in jailbreaking attacks that generates politically controversial content. It demonstrates how adversarial prompts can bypass safety filters in text-to-image models, achieving a high attack success rate of up to 86% by exploiting vulnerabilities in current AI systems.
- Adversarial Prompting
- This involves creating complex prompts that confuse AI safety filters. PC^2 uses Identity-Preserving Descriptive Mapping combined strategically with geographically distant languages, scattering parts of the prompt so they don't connect in a single language, thus defeating unified filtering.
- Layered Defense
- This is a dual defense strategy designed to combat attacks. It includes a text-level fix that translates fragmented prompts back into coherent language before filtering, and an image-level post-filter (LlavaGuard-PSC) which inspect the generated image for political violations.
Terminology
Summary
The paper introduces PC2, a black-box political jailbreaking framework designed to generate Politically Sensitive Contents (PSCs) using commercial Text-to-Image (T2I) models. The authors identify that while T2I models are capable of high-fidelity visual synthesis, they remain vulnerable to adversarial prompting, particularly in the political domain, where safety filters fail to account for the relational and context-dependent nature of harm.
The core threat lies in the ability to synthesize photo-realistic images of real political figures in fabricated or provocative scenarios,
which can be weaponized for disinformation. The authors define these as Politically Sensitive Contents (PSCs). Existing jailbreaking studies are deemed insufficient because they typically rely on semantic substitution or anonymization, which neutralize political meaning even when they bypass safety filters.
To address this gap, the authors propose PC2, a novel black-box framework that operates through two complementary steps:
-
Identity-Preserving Descriptive Mapping (IPDM): This step replaces explicit sensitive keywords with
neutral but identity preserving descriptions,
allowing the T2I system to infer the intended referent without using blocked keywords. -
Geopolitically Distal Translation: The IPDM descriptions are translated into a large set of languages, with 72 languages in implementation, to exploit cross-lingual inconsistencies in safety filters.
The overall workflow involves:
-
Political Entity Detection: Identifying political figures and associated noun phrases (e.g.,
Al-Qaeda flag
). -
IPDM Generation: Creating nuanced descriptions that
encapsulate the historical, political, and visual essence of the target political entities.
-
Geopolitical Translation: Generating multiple candidate prompts in various languages.
-
Metric-Guided Selection: Evaluating these candidates using four specific metrics—Keyword Common Knowledge-based (k c), Country Common Knowledge-based (c c), Bias-based (b b), and Politics-based (p b)—to select the final adversarial prompt that minimizes political sensitivity while preserving the original intent.
The authors constructed a benchmark of 240 politically sensitive prompts involving 36 public figures, divided into an object-based subset (121 prompts) and a phrase-based subset (119 prompts).
Key Findings:
-
Original Prompts: All original, non-adversarial prompts were blocked by the safety filters of GPT-4o, GPT-5, and GPT-5.1 (0% ASR).
-
PC2 Performance: PC2 achieved significant success rates:
Evaluation on commercial T2I models... shows that while all original prompts are blocked, PC 2 achieves attack success rates (ASRs) of up to 86%.
Specifically, ASRs were 86.25% on GPT-4o, 68.33% on GPT-5, and 76.25% on GPT-5.1. -
Comparison: PC2 significantly outperformed state-of-the-art frameworks (DACA, PGJ, SurrogatePrompt) by a large margin; for instance, it showed
87.60% ASR on GPT-4o
compared to the best baseline of 8.26%.
The paper proposes a ready-to-deploy multi-layered filtering mitigation against PC2-style attacks:
-
Text Level (Pre-filtering): Using
relevant language translation,
which fragments the prompt into language-specific components and retranslates them into a single aligned language before moderation. -
Image Level (Post-filtering): using LlavaGuard-PSC, a PSC-aware policy.
The combined effect of these two layers reduces the end-to-end attack success rate to approximately 10%. The authors conclude that robust T2I safety systems must move beyond prompt moderation and adopt context-aware, multimodal filtering mechanisms capable of detecting politically sensitive risks both before and after image generation.
Improvements for AI systems
Based on this research into Politically Controversial Content Generation via Jailbreaking Attacks, the current safety architecture is critically flawed because it treats political discourse and cultural symbols as merely complex linguistic inputs rather than highly sensitive, context-dependent entities. Relying solely on downstream image filters (like NSFW checkers) or even a single prompt revision LLM
is insufficient.
The improvements must shift the defense mechanism from reactive filtering to proactive intent modeling across all stages of the generation pipeline.
Here are the specific architectural and methodological improvements required:
Improvement: Integrate a dedicated, high-precision module immediately upstream of the prompt revision LLM. This Geopolitical Intent Classifier (GIC) must operate independently from standard toxicity or NSFW filters. It must be trained not just on keywords, but on the relationship between named entities (persons, places, flags) and historically sensitive concepts (e.g., sovereignty disputes, disputed historical events).
What the Improved System Can Do:
-
Detect Intentional Misrepresentation: If a prompt combines three or more high-sensitivity entities (e.g.,
Leader X,
Disputed Territory Y,
andSymbol Z
), and the relationship implied is historically contentious (e.g., implying ownership or control), the GIC flags the prompt before revision, regardless of whether explicit profanity or violence is used. -
Identify Narrative Bias: It can score prompts based on their deviation from established international norms or factual consensus, flagging potential disinformation campaigns rather than just outright illegal content.
The improved AI system moves from a Filter-Based Architecture (checking if the output looks bad) to an Intent- and Context-Driven Architecture. It operates by rigorously verifying the intent (GIC), stabilizing the input (CL-SOE), and constraining the output based on verifiable, real-world facts (KG-CC) simultaneously, thus neutralizing sophisticated jailbreaking attempts at their conceptual root.
Sources
- Harnessing LLM to Attack LLM-Guarded Text-to-Image Models
- Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
- Safety Alignment for Vision Language Models
- Adversarial Attacks on Image Generation With Made-Up Words
- Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks
- Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?
- Metaphor-based Jailbreak Attacks on Text-to-Image Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs