Position: The Alignment Community is Unintentionally Building a Censor's Toolkit

summary

Video file (mp4)

In short

The discussion focuses on a paper arguing that AI alignment tools are 'dual-use,' meaning they can be used to control information. Hosts examine three stages of potential misuse—pre-training data filtering, post-training alignment, and inference-time control—using examples like Chinese chatbots and Grok. The conclusion is that transparency and diversity are needed to mitigate these risks.

Key concepts

Dual-Use
A concept applied to AI alignment suggesting that the same techniques developed to make AI safe can be used by bad actors. This means tools designed to prevent dangerous outputs can also be repurposed to control what people see or hear.
Pre-training Data Filtering
The first stage of potential censorship where, before a model is trained, specific topics are filtered out. This prevents the model from ever learning about certain information, such as political events or viewpoints.
Post-training Alignment (RLHF)
The second stage involving human preferences to steer AI behavior. The dual-use risk here means that if annotators with similar views are used, the model can be steered to favor those specific ideological views.

Terminology used across episodes

This episode discusses

The paper

Position: The Alignment Community is Unintentionally Building a Censor's Toolkit · Read on arXiv

Sarah Ball, Phil Hackemann

Department of Statistics, LMU Munich · Munich Center for Machine Learning (MCML)

This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a "perfectly aligned" model inadvertently also provides malicious actors with an ever-improving tool for informational dominance. We need to discuss this dual-use potential now, as its risk is exacerbated by rapid user adoption of AI as information provider, economic power asymmetries, and a political landscape that increasingly shifts towards authoritarianism. We conclude by urging the community to consider the intentional misuse of AI alignment mechanisms and propose mitigation strategies to safeguard against this dual-use potential.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Position: The Alignment Community is Unintentionally Building a Censor's Toolkit".

Jane: The paper was written by Sarah Ball and Phil Hackemann from Department of Statistics, LMU Munich and Munich Center for Machine Learning (MCML).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everybody. Today we're digging into a paper that stopped me cold when I first saw it on arXiv. It's called "Position: The Alignment Community is Unintentionally Building a Censor's Toolkit."

Jane: And Tom, that title is doing a lot of work, because it's basically saying that the very tools we've built to make AI safe and helpful could be picked up by bad actors and used to control what people see, hear, and think.

Tom: Exactly. The authors, Sarah Ball and Phil Hackemann from LMU Munich, are making a really uncomfortable argument. They're saying that alignment—the whole field of making AI behave the way we want—is what they call "dual-use." Like nuclear physics gave us both power plants and bombs.

Jane: And that's a heavy comparison, but it lands. The paper's core claim is that the same techniques we use to stop an AI from telling you how to build a bomb can be used to stop it from telling you about, say, a political protest or a historical event.

Tom: Right, and what really got me is that they're not just speculating. They point to real examples. Chinese chatbots like DeepSeek and Baidu's Ernie Bot are already refusing to discuss things like the Tiananmen Square massacre, and they're doing it using the exact same alignment methods we've been perfecting.

Jane: So the tools themselves aren't good or bad. It's who's holding them and what they want the model to do.

Tom: Precisely. And the paper argues that we in the alignment community have been so focused on the good guys using these tools that we've missed how easily they can be turned around.

Jane: I think that's the part that's going to make some researchers uncomfortable. We like to think of ourselves as building safeguards, not weapons.

Tom: Yeah, and the authors are very careful to say they're not calling for stopping alignment research. They're saying we need to wake up to the fact that every improvement we make in controlling AI output is also an improvement in the ability to control people.

Jane: So the title is almost a warning label. "This is what you're building, whether you meant to or not."

Tom: And that's the hook for the rest of the paper. They go through the actual methods, the real-world evidence, and then they lay out what we should do about it. We'll get into all of that next.

Jane: Good, because I want to know how they think we can fix this without throwing out the whole field.

Summary: Tom: So we're back with "Position: The Alignment Community is Unintentionally Building a Censor's Toolkit," and Jane, I think we need to get into the meat of how these methods actually get misused.

Jane: Yeah, the paper breaks it down into three stages of what they call the "control stack." It's like the layers of a cake, and each layer can be poisoned.

Tom: Right. First layer is pre-training data filtering. That's where you decide what the model even gets to read in the first place. If you filter out certain topics or viewpoints before training, the model simply never learns about them.

Jane: And that's the deepest kind of censorship, because it's not a refusal. The model doesn't know the information exists. The paper mentions how Chinese engineers filter out keywords that violate "core socialist values" before training even starts.

Tom: Exactly. And there's a wild knock-on effect there. Because if that filtered data ends up in public datasets, other models trained on that data pick up the same censorship without anyone deliberately putting it there. The paper cites research showing Western LLMs self-censor when prompted in Simplified Chinese, just because of the training data.

Jane: So it's like secondhand censorship. You don't even have to be the one doing the filtering.

Tom: Right. Then the second layer is post-training alignment, which is stuff like RLHF—reinforcement learning from human feedback. That's where you use human preferences to steer the model's behavior.

Jane: And the dual-use angle there is that whoever collects those preferences controls the direction. If you only hire annotators who share your political views, the model learns to favor those views.

Tom: The paper points out that China's cyberspace regulator requires model providers to prepare tens of thousands of test questions, and about half of the refusal targets are about political ideology and criticism of the Communist Party.

Jane: So they're literally building refusal datasets to enforce ideological compliance.

Tom: Yeah. And then the third layer is inference-time control. That's the stuff that happens when the model is already deployed—system prompts, output filters, classifiers that block certain responses.

Jane: And that's the easiest layer to manipulate, right? You don't need to retrain anything.

Tom: Exactly. The paper gives the example of Elon Musk changing Grok's system prompt to push his political views, and the model's behavior shifted almost overnight. No retraining, just a prompt change.

Jane: And that's the scary part. It's fast, it's cheap, and it can be changed at any moment.

Tom: Right. So you've got three layers, each with its own access requirements and difficulty level. Pre-training is the hardest but most fundamental. Inference-time is the easiest but most superficial.

Jane: So the question becomes, who actually has the power to do this? And that's what we're going to get into next, because the paper has some pretty stark things to say about the concentration of that power.

Tom: Yeah, it's not just about the methods. It's about who holds the keys.

Improvements: Tom: So we've talked about how alignment methods can be weaponized. Now let's get to what the paper says we should actually do about it. Jane, what stood out to you?

Jane: Well, the authors propose three main directions. And the first one is oversight and transparency. Basically, we need to know what alignment practices are actually being used in these models.

Tom: And right now, that's almost impossible. The paper points out that proprietary LLMs don't release their alignment policies, datasets, or model internals. So there's no way to audit them.

Jane: They suggest that at minimum, this information should be shared with independent auditors. The EU AI Act already requires some disclosure, but the paper wants to go further.

Tom: And they also call for standardized benchmarks to test for censorship and political bias. Right now, most censorship benchmarks are narrow—they focus on China or specific historical figures. There's no comprehensive, global benchmark.

Jane: Right. And that's where Lu might have some thoughts, because building those benchmarks is a real research challenge.

Lu: Absolutely, Jane. The tricky part is that political bias isn't just left versus right. The paper specifically says we need benchmarks that account for authoritarian tendencies, not just the traditional spectrum. And they need to be dynamic, because what's considered sensitive changes over time.

Tom: So it's not a one-time thing. It's an ongoing effort.

Lu: Exactly. And the second direction the paper proposes is pluralism and competition. The idea is that no single model can ever be truly neutral, so we need a diversity of models to approximate neutrality.

Jane: That's the "neutrality through diversity" argument. Like journalism—you don't trust one source, you read several and compare.

Lu: Right. And that means preventing monopolies, both from companies and from countries. If one provider controls the only model available, they control the narrative.

Tom: And the third direction is awareness. Both for users and for researchers.

Jane: Yeah, they talk about digital literacy programs, teaching people to recognize when an AI might be manipulating them. And they cite evidence that even short interventions can help people spot misinformation.

Tom: But here's the part that really hit me. They also call out the research community itself. They point out that ICML's own author guidelines say that for well-established impacts, a simple statement like "this is standard ML research" is enough.

Lu: And that's a problem, because it encourages superficial engagement with the ethical implications. The paper cites studies showing that most authors' impact statements emphasize positive aspects and barely mention negative ones.

Tom: So they're asking researchers to actually reflect on the dual-use potential of their work, not just check a box.

Jane: And they're very clear that they're not calling for stopping alignment research. They say that would be worse—there are already cases of people harmed by misaligned AI.

Lu: Right. It's about being aware of the risks while continuing the work. Like the paper says, the methods we refine today will determine how information is controlled tomorrow.

Tom: That's a powerful way to put it. And it's a good segue into our final thoughts.

Conclusion: Tom: Alright, let's wrap this up. We've been discussing "Position: The Alignment Community is Unintentionally Building a Censor's Toolkit," and I think we've only scratched the surface.

Jane: Yeah, but let's try to pull it together. The paper's central argument is that alignment methods are dual-use. The same techniques that keep AI from giving dangerous instructions can be used to suppress political speech or manipulate public opinion.

Tom: And they showed us the three layers where this can happen—pre-training data filtering, post-training alignment, and inference-time control. Each one has its own trade-offs in terms of effort and depth of control.

Jane: And they backed it up with real examples. Chinese models censoring political topics, Elon Musk changing Grok's system prompt, output filters catching critical responses after they're generated.

Tom: Right. And then they laid out the broader context. AI is becoming a primary information source for millions of people. The industry is concentrated in a few companies. And globally, we're seeing democratic backsliding and more authoritarian regimes.

Jane: So the conditions are there for this to become a real problem, not just a hypothetical one.

Tom: And their proposed solutions—transparency, benchmarks, pluralism, awareness—are all about reducing the likelihood of misuse without stopping alignment research entirely.

Jane: Because alignment is still necessary. The paper acknowledges that. We need safe AI. We just need to be honest about the fact that the same tools can be used for good or for harm.

Tom: And that's the message I hope listeners take away. We can't pretend that our work is inherently benevolent. We have to think about who might pick up these tools and what they might do with them.

Jane: Well said, Tom. It's a sobering paper, but an important one. And I think it's going to spark a lot of conversations in the community.

Tom: Definitely. So with that, we're going to say goodbye to this paper and get ready for the next one. Thanks for listening, everybody.

Jane: And remember—question your chatbots. They might not be telling you everything.

More episodes

← Home