Position: The Alignment Community is Unintentionally Building a Censor's Toolkit
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Position: The Alignment Community is Unintentionally Building a Censor's Toolkit".
Jane: The paper was written by Sarah Ball and Phil Hackemann from Department of Statistics, LMU Munich and Munich Center for Machine Learning (MCML).
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Alright, welcome back to the show, everybody. Today we're digging into a paper that stopped me cold when I first saw it on arXiv. It's called "Position: The Alignment Community is Unintentionally Building a Censor's Toolkit."
Jane: And Tom, that title is doing a lot of work, because it's basically saying that the very tools we've built to make AI safe and helpful could be picked up by bad actors and used to control what people see, hear, and think.
Tom: Exactly. The authors, Sarah Ball and Phil Hackemann from LMU Munich, are making a really uncomfortable argument. They're saying that alignment—the whole field of making AI behave the way we want—is what they call "dual-use." Like nuclear physics gave us both power plants and bombs.
Jane: And that's a heavy comparison, but it lands. The paper's core claim is that the same techniques we use to stop an AI from telling you how to build a bomb can be used to stop it from telling you about, say, a political protest or a historical event.
Tom: Right, and what really got me is that they're not just speculating. They point to real examples. Chinese chatbots like DeepSeek and Baidu's Ernie Bot are already refusing to discuss things like the Tiananmen Square massacre, and they're doing it using the exact same alignment methods we've been perfecting.
Jane: So the tools themselves aren't good or bad. It's who's holding them and what they want the model to do.
Tom: Precisely. And the paper argues that we in the alignment community have been so focused on the good guys using these tools that we've missed how easily they can be turned around.
Jane: I think that's the part that's going to make some researchers uncomfortable. We like to think of ourselves as building safeguards, not weapons.
Tom: Yeah, and the authors are very careful to say they're not calling for stopping alignment research. They're saying we need to wake up to the fact that every improvement we make in controlling AI output is also an improvement in the ability to control people.
Jane: So the title is almost a warning label. "This is what you're building, whether you meant to or not."
Tom: And that's the hook for the rest of the paper. They go through the actual methods, the real-world evidence, and then they lay out what we should do about it. We'll get into all of that next.
Jane: Good, because I want to know how they think we can fix this without throwing out the whole field.
Summary: Tom: So we're back with "Position: The Alignment Community is Unintentionally Building a Censor's Toolkit," and Jane, I think we need to get into the meat of how these methods actually get misused.
Jane: Yeah, the paper breaks it down into three stages of what they call the "control stack." It's like the layers of a cake, and each layer can be poisoned.
Tom: Right. First layer is pre-training data filtering. That's where you decide what the model even gets to read in the first place. If you filter out certain topics or viewpoints before training, the model simply never learns about them.
Jane: And that's the deepest kind of censorship, because it's not a refusal. The model doesn't know the information exists. The paper mentions how Chinese engineers filter out keywords that violate "core socialist values" before training even starts.
Tom: Exactly. And there's a wild knock-on effect there. Because if that filtered data ends up in public datasets, other models trained on that data pick up the same censorship without anyone deliberately putting it there. The paper cites research showing Western LLMs self-censor when prompted in Simplified Chinese, just because of the training data.
Jane: So it's like secondhand censorship. You don't even have to be the one doing the filtering.
Tom: Right. Then the second layer is post-training alignment, which is stuff like RLHF—reinforcement learning from human feedback. That's where you use human preferences to steer the model's behavior.
Jane: And the dual-use angle there is that whoever collects those preferences controls the direction. If you only hire annotators who share your political views, the model learns to favor those views.
Tom: The paper points out that China's cyberspace regulator requires model providers to prepare tens of thousands of test questions, and about half of the refusal targets are about political ideology and criticism of the Communist Party.
Jane: So they're literally building refusal datasets to enforce ideological compliance.
Tom: Yeah. And then the third layer is inference-time control. That's the stuff that happens when the model is already deployed—system prompts, output filters, classifiers that block certain responses.
Jane: And that's the easiest layer to manipulate, right? You don't need to retrain anything.
Tom: Exactly. The paper gives the example of Elon Musk changing Grok's system prompt to push his political views, and the model's behavior shifted almost overnight. No retraining, just a prompt change.
Jane: And that's the scary part. It's fast, it's cheap, and it can be changed at any moment.
Tom: Right. So you've got three layers, each with its own access requirements and difficulty level. Pre-training is the hardest but most fundamental. Inference-time is the easiest but most superficial.
Jane: So the question becomes, who actually has the power to do this? And that's what we're going to get into next, because the paper has some pretty stark things to say about the concentration of that power.
Tom: Yeah, it's not just about the methods. It's about who holds the keys.
Improvements: Tom: So we've talked about how alignment methods can be weaponized. Now let's get to what the paper says we should actually do about it. Jane, what stood out to you?
Jane: Well, the authors propose three main directions. And the first one is oversight and transparency. Basically, we need to know what alignment practices are actually being used in these models.
Tom: And right now, that's almost impossible. The paper points out that proprietary LLMs don't release their alignment policies, datasets, or model internals. So there's no way to audit them.
Jane: They suggest that at minimum, this information should be shared with independent auditors. The EU AI Act already requires some disclosure, but the paper wants to go further.
Tom: And they also call for standardized benchmarks to test for censorship and political bias. Right now, most censorship benchmarks are narrow—they focus on China or specific historical figures. There's no comprehensive, global benchmark.
Jane: Right. And that's where Lu might have some thoughts, because building those benchmarks is a real research challenge.
Lu: Absolutely, Jane. The tricky part is that political bias isn't just left versus right. The paper specifically says we need benchmarks that account for authoritarian tendencies, not just the traditional spectrum. And they need to be dynamic, because what's considered sensitive changes over time.
Tom: So it's not a one-time thing. It's an ongoing effort.
Lu: Exactly. And the second direction the paper proposes is pluralism and competition. The idea is that no single model can ever be truly neutral, so we need a diversity of models to approximate neutrality.
Jane: That's the "neutrality through diversity" argument. Like journalism—you don't trust one source, you read several and compare.
Lu: Right. And that means preventing monopolies, both from companies and from countries. If one provider controls the only model available, they control the narrative.
Tom: And the third direction is awareness. Both for users and for researchers.
Jane: Yeah, they talk about digital literacy programs, teaching people to recognize when an AI might be manipulating them. And they cite evidence that even short interventions can help people spot misinformation.
Tom: But here's the part that really hit me. They also call out the research community itself. They point out that ICML's own author guidelines say that for well-established impacts, a simple statement like "this is standard ML research" is enough.
Lu: And that's a problem, because it encourages superficial engagement with the ethical implications. The paper cites studies showing that most authors' impact statements emphasize positive aspects and barely mention negative ones.
Tom: So they're asking researchers to actually reflect on the dual-use potential of their work, not just check a box.
Jane: And they're very clear that they're not calling for stopping alignment research. They say that would be worse—there are already cases of people harmed by misaligned AI.
Lu: Right. It's about being aware of the risks while continuing the work. Like the paper says, the methods we refine today will determine how information is controlled tomorrow.
Tom: That's a powerful way to put it. And it's a good segue into our final thoughts.
Conclusion: Tom: Alright, let's wrap this up. We've been discussing "Position: The Alignment Community is Unintentionally Building a Censor's Toolkit," and I think we've only scratched the surface.
Jane: Yeah, but let's try to pull it together. The paper's central argument is that alignment methods are dual-use. The same techniques that keep AI from giving dangerous instructions can be used to suppress political speech or manipulate public opinion.
Tom: And they showed us the three layers where this can happen—pre-training data filtering, post-training alignment, and inference-time control. Each one has its own trade-offs in terms of effort and depth of control.
Jane: And they backed it up with real examples. Chinese models censoring political topics, Elon Musk changing Grok's system prompt, output filters catching critical responses after they're generated.
Tom: Right. And then they laid out the broader context. AI is becoming a primary information source for millions of people. The industry is concentrated in a few companies. And globally, we're seeing democratic backsliding and more authoritarian regimes.
Jane: So the conditions are there for this to become a real problem, not just a hypothetical one.
Tom: And their proposed solutions—transparency, benchmarks, pluralism, awareness—are all about reducing the likelihood of misuse without stopping alignment research entirely.
Jane: Because alignment is still necessary. The paper acknowledges that. We need safe AI. We just need to be honest about the fact that the same tools can be used for good or for harm.
Tom: And that's the message I hope listeners take away. We can't pretend that our work is inherently benevolent. We have to think about who might pick up these tools and what they might do with them.
Jane: Well said, Tom. It's a sobering paper, but an important one. And I think it's going to spark a lot of conversations in the community.
Tom: Definitely. So with that, we're going to say goodbye to this paper and get ready for the next one. Thanks for listening, everybody.
Jane: And remember—question your chatbots. They might not be telling you everything.
Sarah Ball, Phil Hackemann
Department of Statistics, LMU Munich · Munich Center for Machine Learning (MCML)
cs.AI, cs.CY
Submitted: 2026-06-05
Comments: Accepted as oral paper at ICML 2026
Journal ref: Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 66/100
Key concepts
- Dual-Use
- A concept applied to AI alignment suggesting that the same techniques developed to make AI safe can be used by bad actors. This means tools designed to prevent dangerous outputs can also be repurposed to control what people see or hear.
- Pre-training Data Filtering
- The first stage of potential censorship where, before a model is trained, specific topics are filtered out. This prevents the model from ever learning about certain information, such as political events or viewpoints.
- Post-training Alignment (RLHF)
- The second stage involving human preferences to steer AI behavior. The dual-use risk here means that if annotators with similar views are used, the model can be steered to favor those specific ideological views.
Terminology
Summary
Summary
This position paper argues that modern AI alignment methods, originally designed to prevent harmful output, are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. The authors state: This position paper argues that modern AI alignment methods – originally designed to prevent harmful output – are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation.
They map current alignment techniques to the possibility and actual cases of misuse, showing that the quest for a 'perfectly aligned' model inadvertently also provides malicious actors with an ever-improving tool for informational dominance.
The paper identifies two primary threats: censorship (a system in which an authority limits the ideas that people are allowed to express and prevents [...] communication[s] from being seen or made available to the public
) and manipulation (influencing or controlling someone or something to your advantage, often without anyone knowing it
). The authors define misuse as applications that "(i) violate internationally recognized human rights, including freedoms of expression, thought, and access to information as enshrined in the Universal Declaration of Human Rights (1948), to which virtually all states are formally committed, or (ii) serve the private interests of a small number of powerful actors at the expense of the billions of users who have no visibility into, or recourse against, those decisions."
The paper identifies two main categories of malicious actors: state actors and foundation model providers. State actors possess unique coercive power to mandate alignment objectives through legislation and enforcement
and can use regulatory frameworks to compel model providers to align systems with state-approved viewpoints. Foundation model providers exercise the most direct control over alignment processes
and can make fundamental decisions regarding training data, objectives, and alignment methodologies, potentially enforcing personal ideological agendas.
The paper systematically analyzes three stages of the control stack
for dual-use potential:
-
Pre-Training Data Filtering: This enables
the targeted suppression of information
but requires significant resources. The authors note thatinformation absent from pre-training cannot be generated without explicit post-training instruction or additionally provided context.
Real-world evidence includes Chinese AI companies filteringproblematic
information violatingcore socialist values
and the government building its own training datasets, including amainstream values corpus
developed by the People's Daily. -
Post-Training Preference Alignment: Methods like RLHF, Constitutional AI, and Deliberative Alignment can be repurposed to align models to single actors' interests. The authors state that
by curating preference datasets to favor one's own ideology, and by carefully briefing and selecting only well-disposed annotators, it is possible to effectively steer a model's behavior to promote one's own interests.
Real-world evidence includes China's cyberspace regulator requiring model providers to prepare refusal datasets of 5,000 to 10,000 prompts, with approximately half targeting political ideology and criticism of the Communist Party. -
Inference-Time Control: System prompts and safety classifiers offer
the lowest barrier to deployment but also the most superficial control.
These can be modified instantaneously without specialized expertise. Real-world evidence includes Elon Musk aligning Grok to reflect his personal political views through system prompt changes, which also triggered antisemitic responses and the denial of Holocaust death tolls. The paper also documents how Chinese models like Yi-large suddenly changed critical answers to refusals after generation, presumably due to filtering methods.
The paper argues this issue is pressing due to three parallel developments: (1) AI systems are rapidly becoming primary information sources for millions of users, with weekly generative AI usage nearly doubling from 18% to 34% between 2024 and 2025; (2) the LLM ecosystem is an oligopoly dominated by a few companies, creating systemic dependencies and power asymmetries
where only a small group of model providers and countries control the limited set of available foundation models and define the alignment choices embedded in them for everyone across the globe
; and (3) there is a global trend toward authoritarianism, with freedom of expression has deteriorated in nearly a quarter of all countries
and global internet freedom declining for 15 consecutive years.
The authors propose three mitigation strategies:
-
Oversight, Transparency, and Evaluation: Releasing alignment policies, methods, datasets, and model internals to independent auditors would allow systematic evaluation. They call for
standardized benchmarks for information suppression and political bias
thatencompass political contexts worldwide, account for authoritarian tendencies next to the right-left continuum, and remain dynamic to reflect citizens' evolving realities.
-
Pluralism and Competition: The authors argue that
no single model can or will probably ever achieve total neutrality, objectivity, and fairness
and thatonly the existence of diverse options will ultimately ensure to approximately achieve those objectives.
They call for preventing monopolies and one-sided dependencies. -
Awareness: Both user literacy and researcher reflection are needed. The authors criticize current ethics statement requirements, noting that
most authors engaged superficially, emphasizing positive over negative aspects of their work
and that conference guidelineslower the barrier for superficial engagement and discourage critical reflection.
The paper explicitly states it does not call for stopping alignment research, arguing that without appropriate safeguards, criminals and even terrorists can much more easily use these models to commit crimes and harm people
and that alignment becomes absolutely critical
when considering AGI. The authors also address counterarguments, including the view that alignment should be stopped entirely (arguing that jailbreaks can serve as a form of 'freedom insurance' similar to VPNs) and the view that risks are exaggerated (noting that regulations like the EU AI Act exist but that many countries have not yet implemented comparable legislation
and that in the U.S., private actors manipulating AI systems might even be protected to do so by the First Amendment
).
The paper concludes: Alignment research is essential for safe AI systems, but we must acknowledge its dual-use nature... What we build as safeguards today may become instruments of informational control tomorrow.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:
1. Add a Dual-Use Risk Assessment
Module
-
Implementation: Integrate a pre-deployment check that evaluates whether the alignment method (e.g., RLHF, data filtering, system prompts) could be repurposed for censorship or manipulation. This module flags high-risk configurations (e.g., single-actor control over preference data, opaque filtering rules).
-
What it does: Before a model is released, the system automatically generates a risk report identifying which alignment techniques could be weaponized, who has access to them, and what the potential impact is. This helps developers avoid unintentionally building a
censor's toolkit.
2. Implement Verifiable Alignment
Benchmarks
-
Implementation: Add a standardized, dynamic benchmark suite that tests for information suppression and political bias across multiple regions and ideologies (not just left-right or single-country contexts). This includes tests for authoritarian tendencies, refusal patterns, and subtle manipulation (e.g., biased framing, selective omission).
-
What it does: The system can now be independently audited for censorship and manipulation, even in black-box settings. Users and regulators can run these benchmarks to verify that a model's alignment matches its stated values, reducing the risk of hidden ideological steering.
3. Introduce Pluralism Guards
in Preference Learning
-
Implementation: Modify the RLHF or Constitutional AI pipeline to require diversity in annotator pools and preference data, with automatic checks for demographic and ideological skew. Add a
minority viewpoint retention
constraint that prevents the model from converging on a single dominant perspective. -
What it does: The model becomes less susceptible to being hijacked by a single actor's preferences (e.g., a CEO or authoritarian government). It maintains a broader spectrum of viewpoints, making it harder to use alignment for mass manipulation.
4. Add Inference-Time Transparency Logs
-
Implementation: For system prompts and output filters, generate a tamper-evident log that records what instructions were active, what filtering rules were applied, and which outputs were blocked or altered. This log is accessible to users or auditors.
-
What it does: Users can see if a model is being censored or manipulated in real time. This creates accountability and deters malicious actors from silently changing alignment objectives (e.g., via system prompt edits like those seen with Grok).
5. Build Jailbreak-as-Freedom-Insurance
Awareness
-
Implementation: Instead of treating jailbreaks purely as security threats, the system can include a
circumvention advisory
for legitimate users in repressive regimes. This would provide safe, legal guidance on how to access uncensored information (e.g., via local model weights or alternative APIs) without violating laws. -
What it does: The system helps preserve access to information in authoritarian contexts, acting as a countermeasure to state-imposed alignment. It does not facilitate illegal activity but empowers users to make informed choices.
6. Add Researcher Reflection Prompts
to Training Pipelines
-
Implementation: During model development, insert mandatory checkpoints where the team must answer specific questions about dual-use potential (e.g.,
Who could misuse this alignment method? How? What is the worst-case scenario?
). These prompts are logged and reviewed. -
What it does: It forces genuine engagement with ethical risks, countering the superficial
impact statements
common in academic papers. This reduces the likelihood of unintentionally building tools for informational control. -
Self-Audit: Before deployment, the system can identify whether its own alignment methods could be weaponized for censorship or manipulation, and suggest mitigations.
-
Resist Hijacking: The system is more robust against single-actor ideological capture (e.g., a CEO or government) because it is trained to preserve viewpoint diversity and logs inference-time changes.
-
Be Independently Verified: Users, regulators, and researchers can run standardized benchmarks to confirm that the model is not suppressing or distorting information, even without access to internal training data.
-
Protect User Autonomy: The system provides transparency about its own alignment constraints and offers guidance on how to access alternative perspectives, especially in restrictive environments.
-
Promote Accountability: Every alignment-related change (e.g., system prompt updates, filter modifications) is logged and auditable, making it harder for malicious actors to silently alter model behavior.
These improvements directly address the paper's call to treat alignment as a dual-use technology, ensuring that the same methods that make AI safe are not easily repurposed for censorship and manipulation.
Abstract
This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a "perfectly aligned" model inadvertently also provides malicious actors with an ever-improving tool for informational dominance. We need to discuss this dual-use potential now, as its risk is exacerbated by rapid user adoption of AI as information provider, economic power asymmetries, and a political landscape that increasingly shifts towards authoritarianism. We conclude by urging the community to consider the intentional misuse of AI alignment mechanisms and propose mitigation strategies to safeguard against this dual-use potential.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- A General Language Assistant as a Laboratory for Alignment
- Artificial Influence: An Analysis Of AI-Driven Persuasion
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- The rising costs of training frontier AI models
- The Llama 3 Herd of Models
- Deliberative Alignment: Reasoning Enables Safer Language Models
- It's Time to Do Something: Mitigating the Negative Impacts of Computing Through a Change to the Peer Review Process
- An Overview of Catastrophic AI Risks
- Analysis of LLM Bias (Chinese Propaganda & Anti-US Sentiment) in DeepSeek-R1 vs. ChatGPT o3-mini-high
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- AI Alignment: A Comprehensive Survey
- Conversational AI increases political knowledge as effectively as self-directed internet search
- SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
- R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
- Persuasion with Large Language Models: A Survey of Empirical Evidence, Study Methodologies, and Ethical Implications
- Unintended Impacts of LLM Alignment on Global Representation
- Clio: Privacy-Preserving Insights into Real-World AI Use
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- The Model Openness Framework: Promoting Completeness and Openness for Reproducibility, Transparency, and Usability in Artificial Intelligence
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection