Ideology-Based LLMs for Content Moderation

arXiv:2510.25805 · cs.CL · Submitted 2025-10-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Ideology-Based LLMs for Content Moderation".

Jane: Persona conditioning introduces subtle ideological biases into LLM outputs, raising concerns about AI systems that may reinforce partisan perspectives under the guise of neutrality.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We started by looking at the title and the people who did this research, so we know exactly what we're diving into with today's topic. The paper is titled "Ideology-Based LLMs for Content Moderation." It’s authored by Stefano Civelli, Pietro Bernardelle, Nardiena Pratama, and Gianluca Demartini from the University of Queensland.

Jane: Those are some heavy hitters in the AI research world, Tom. So what's the main point we should get out of that title? It suggests they aren't just looking at how good a model is at spotting bad content, but how it behaves when it has an assigned political identity attached to it.

Tom: That’s right. Essentially, the core idea is examining how adding these persona layers influences the consistency and fairness of harmful content classification across different types of AI architectures and even different kinds of data like text versus vision.

Lu: From a creative standpoint, this opens up so many possibilities for building more nuanced moderation systems where the context isn't just about words, but about the underlying societal lens we program into the model.

Meng: I'm interested in what they found regarding those different architectures and modalities; does that mean some models are way more susceptible to these persona shifts than others?

Lalam: I think it’s really interesting how this approach might help us understand how to build AI that can actually reflect the complexity of human society in its moderation choices.

The paper's summary: Tom: Now we move into the actual summary of what this study found regarding these persona-conditioned models. The researchers used a dataset with two hundred thousand synthetically generated personas and mapped them onto a political compass to select some really extreme viewpoints for testing <ref:2510.25805#pg0>.

Jane: So, what did that process reveal about how these extreme personas actually interact when the AI is doing the classification task? It’s not just about which model performs best overall, but how they behave relative to each other.

Tom: Well, while the headline numbers showed that persona conditioning didn't have a massive impact on general accuracy across all models tested, a closer look at their behavior revealed some very specific shifts in how the models operated.

Lu: The paper highlighted that models tend to align more closely with personas from the same political ideology, which strengthens consistency within those ideological groups but widens the gap between different ideological clusters.

Meng: That suggests that if we want consistent moderation, we might need to be very careful about how we define or apply those personas in our system design.

Lalam: It makes sense; if the model is locked into one viewpoint, it’s going to be much more predictable within its own ideological bubble than when it's trying to balance different viewpoints.

The paper's improvements: Tom: Moving on, we need to talk about what the authors suggest as potential improvements for this work and what they think those changes could do for future research. They pointed toward a few areas where the methodology itself could be refined to get even deeper insights into this phenomenon.

Jane: So, if we look at how they suggest tweaking their approach, are they focusing on fixing the measurement or changing how the personas are selected? What are these suggested improvements for the study?

Tom: They suggest improving content moderation by implementing something called a "Persona-Aware Bias Detection Layer." This layer would specifically check not just for toxicity, but for ideological consistency by seeing if a model's judgment deviates from the expected agreement patterns.

Lu: That sounds like a really practical step toward creating automated checks that look beyond simple binary classification and start looking at the underlying bias structure.

Meng: From an engineering standpoint, that layer could be integrated into deployment pipelines to catch these shifts in real-time rather than just post-hoc analysis of the results.

Lalam: I think if we can build a system that flags when the model's behavior starts pulling away from its expected ideological alignment, it could give us much better control over how we deploy these systems.

Conclusion: Tom: So, let's wrap up this deep dive into "Ideology-Based LLMs for Content Moderation." The main takeaway is that persona conditioning acts as a strong vector for introducing and amplifying ideological biases, even if it doesn't hurt the basic performance of the AI.

Jane: It really underscores how subtle these shifts in behavior can be when we assign an identity to a model, showing us that we need to pay close attention to those internal interpretive biases.

Tom: Exactly. The study shows that while models might seem neutral on the surface, they develop partisan asymmetries in moderation, especially as the models get bigger and more capable.

Lu: The implication is that we have to be extremely careful about ensuring fairness in these systems, paying attention to the subtle ways LLMs interpret and embody the identities we assign them, and the ideological biases that can emerge as a result.

Meng: I agree with Lu; from an engineering perspective, this means testing across different model scales and modalities is essential to see how this divergence scales up as capability increases.

Lalam: And for us in development, it means we need to design ways to actively counteract these persona-induced biases before they become deeply embedded in the system's core function.

Tom: That’s all the time we have for today on this paper. We'll be keeping an eye out for more research coming across arXiv next week.

The University of Queensland

cs.CL

Submitted: 2025-10-29

Updated: 2025-10-29

Code: https://github.com/facebookresearch/fine_grained_hateful_memes2https:

Importance score: 90/100

The gist: Persona conditioning introduces subtle ideological biases into LLM outputs, raising concerns about AI systems that may reinforce partisan perspectives under the guise of neutrality.

Key concepts

Persona Conditioning
This involves assigning a specific political identity or set of beliefs to an LLM through a description. The research found this acts as a powerful tool that introduces and amplifies inherent ideological biases into the model's behavior, even when the goal is neutrality.
Political Compass Test (PCT)
This is a method used to map various political viewpoints onto a two-dimensional scale. Researchers used this test to categorize 200,000 synthetic personas across six different language models, allowing them to select 'extreme' ideological examples for testing.
Partisan Asymmetry
This refers to the observed differences in how LLMs moderate content based on the political leanings of the persona they are adopting. Larger models showed this effect strongly, where left personas prioritized protecting anti-left speech while right personas did the opposite.
Intra- and Inter-Group Agreement
This measures how much different personas agree with each other. The study found that agreement within the same political group (intra-ideology) was consistently higher than agreement between different groups (inter-ideology), suggesting ideological cohesion is stronger.

Terminology

Summary

Persona conditioning introduces subtle ideological biases into LLM outputs, raising concerns about AI systems that may reinforce partisan perspectives under the guise of neutrality.

Overview and Research Questions

This study examines how persona adoption influences the consistency and fairness of harmful content classification across different LLM architectures, model sizes, and content modalities (language vs. vision). The research addresses four specific questions: RQ1 asks how political ideology encoded in persona descriptions affects decision consistency; RQ2 investigates whether personas with different ideological leanings systematically differ in their propensity to label content as harmful; RQ3 explores the extent to which moderation behavior aligns with shared political ideologies, examining intra- and inter-group agreement patterns; and RQ4 assesses susceptibility of different LLM architectures, model sizes, or modalities to persona-induced behavioral divergence.

Methodology: Persona Mapping and Selection

The researchers utilized the PersonaHub dataset, containing 200,000 synthetically generated persona descriptions. These descriptions were mapped onto a two-dimensional political compass using the Political Compass Test (PCT), resulting in a distribution of political coordinates for each persona across six different language models. From these distributions, extreme personas were selected using two strategies:

  1. Corner Selection: Selecting 100 personas from each of the four compass quadrants to maximize both ideological extremity and internal consistency.

  2. Economic-Axis Selection: Selecting 200 personas from the far economic left and 200 from the far right to isolate the effect of economic polarization.

Content Moderation Evaluation

The selected extreme personas were then used to prompt the same language models to classify a fixed set of harmful content examples drawn from three datasets: Hate-Identity (text), Facebook Hateful Memes (vision), and Contextual Abuse Dataset (text). The evaluation involved measuring overall moderation capabilities, detection sensitivity to generic harmful content, agreement patterns by ideology, and partisan asymmetries in political hate speech moderation.

Key Findings on Performance and Bias

Headline performance metrics suggested personas had little impact on overall classification accuracy when aggregated across all personas. However, a closer analysis revealed systematic behavioral shifts:

  1. Consistency within Ideology: Models tend to align more closely with personas from the same political ideology, strengthening within-ideology consistency while widening divergence across ideological groups.

  2. Sensitivity Variation: Personas with different ideological leanings display distinct propensities to label content as harmful, indicating that the lens through which a model interprets input can subtly shape its judgments.

  3. Agreement Patterns: Agreement analyses showed that intra-ideology agreement is systematically higher than inter-ideology agreement, and this cohesion grows stronger with model scale.

Partisan Asymmetries in Political Hate Speech

When moderating politically targeted content, the study found marked partisan asymmetries. For smaller models, personas primarily reflected a global sensitivity shift along the political spectrum. However, larger models exhibited a more complex pattern: "left personas show heightened sensitivity to anti-left hate speech (OR > 1), while right personas display the reverse, becoming more sensitive to anti-right hate speech (OR < 1). This reversal suggests that ideological alignment conditions the model to prioritize protection of its “in-group” while downplaying harm directed at opponents. The strength of this effect appears to scale with model size."

Conclusion and Implications

The findings suggest that persona conditioning is a powerful vector for introducing and amplifying ideological biases. While it does not undermine baseline competence, it introduces systematic shifts in behavior that manifest as partisan asymmetries in moderation. This raises the risk that AI models may inadvertently reinforce partisan viewpoints while presenting themselves as neutral arbiters, especially as models grow more capable. The study concludes that ensuring fairness requires attention to the subtle ways in which LLMs interpret and embody the identities we assign them, and the ideological biases that can emerge as a result.

Limitations

The researchers acknowledged limitations, including the use of synthetic personas from PersonaHub which may not fully capture real-world nuance, the simplified nature of the Political Compass Test, and restrictions to a limited set of open-source models. Future work is suggested to include models with extensive safety fine-tuning and develop methods to counteract persona-induced bias. The study ultimately highlights that persona prompting is not a neutral interface for customization but a powerful vector for introducing and amplifying ideological biases.

References

[1] Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023. Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis 31, 3 (Feb. 2023), 337–351.

[2] Pietro Bernardelle, Stefano Civelli, Leon Fröhling, Riccardo Lunardi, Kevin Roitero, and Gianluca Demartini. 2025. Political Ideology Shifts in Large Language Models.

Improvements for AI systems

Here are specific improvements for AI systems based on the findings of Ideology-Based LLMs for Content Moderation:

  1. Improve content moderation by implementing a Persona-Aware Bias Detection Layer. This layer would analyze LLM outputs not just for toxicity, but for ideological consistency. It would use the established political compass coordinates derived from persona conditioning to flag when a model's judgment (e.g., labeling text as hate speech) deviates significantly from its expected intra-ideological agreement pattern, or when it exhibits the partisan asymmetry observed in larger models (where left personas are more trigger-happy and right personas show defensive bias).

  2. Develop adaptive moderation thresholds based on input context and inferred persona influence. Instead of a static safety filter, the system should dynamically adjust its sensitivity (precision vs. recall) based on whether the content being moderated aligns with content typically associated with specific ideological personas (e.g., if a user profile suggests a right-leaning persona, the model applies the conservative threshold seen in TR personas; if it suggests left-leaning, it applies the more sensitive threshold seen in BL personas).

  3. Enhance robustness against adversarial prompting by training models using Ideological In-Group Defense reinforcement learning (RLHF). This involves specifically rewarding models when they maintain high agreement within their assigned ideological cluster while penalizing them for exhibiting overt partisan asymmetry or failing to recognize the defensive bias observed in larger models when moderating politically charged content.

  4. Implement a multi-modal confidence scoring system that accounts for modality-specific instability. For vision-language models (VLMs), the system must incorporate a lower confidence score when classifying hate speech if the input is complex or ambiguous, reflecting the observed lower inter-rater reliability and greater instability in VLM outputs compared to text-only models.

  5. Improve fairness auditing by mandating model evaluation across different scales and modalities. Future deployment should require testing not just against baseline performance, but explicitly measuring how agreement patterns (intra vs. inter-ideology) scale with model size (e.g., 7B vs. 70B parameters) and modality, ensuring that the system does not amplify ideological biases as it scales up in capability.

Sources

Related papers