Large Language Models are often politically extreme, usually ideologically inconsistent, and persuasive even in informational contexts
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Large Language Models are often politically extreme, usually ideologically inconsistent, and persuasive even in informational contexts".
Jane: This paper challenges the prevailing academic consensus that Large Language Models (LLMs) exhibit only small, moderate, or slightly left-leaning political biases.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, what's the core of this study? The authors are comparing thirty-one different LLMs against three very different real groups: U.S. federal legislators, Supreme Court Justices, and a representative sample of U.S. voters to see how they align ideologically.
Jane: They are essentially saying that the consensus among researchers about LLMs having only small biases is flawed because they’re using established methods to measure ideology, like those used in political science, to test these models.
Lu: The paper finds that even though LLMs might look moderate on the surface when averaged out across all topics, they are often quite extreme on specific dimensions of ideology.
Meng: Extreme how? Give us an example so we can gauge how much this actually matters in the real world, not just theoretical scores.
Lalam: The abstract points out that these LLMs are not just exhibiting small biases; they show a net result of offsetting extremes when compared to political groups.
Tom: Exactly! For instance, when looking at national partisan politics, the LLMs express moderate preferences on average, but then on the second dimension of ideology, they are much more liberal overall.
Jane: That’s a nuanced finding; it means we can't just look at one single score and assume the model is stable across every issue, which is something I find really important to understand simply.
Lu: The comparison against the Supreme Court Justices also shows that most LLMs fall somewhere in the middle of the liberal and conservative blocs, but they aren't perfectly balanced.
Meng: That suggests that even when dealing with highly deliberative groups like judges, there’s still significant ideological variation across the models themselves.
Lalam: I think this diversity in extremity is key because it shows the underlying structure of these models isn't monolithic; they hold a mix of counterbalancing extreme opinions.
The paper's summary: Tom: Moving on from the comparisons, the authors show that LLMs aren't just moderately biased; they are ideologically inconsistent in a way that mirrors how real voters operate.
Jane: That’s a really insightful point, Tom. It suggests that instead of being reliably left or right, an LLM can be very liberal on healthcare but strongly conservative on gun control simultaneously.
Lu: The paper quantifies this inconsistency by showing that models hold these kinds of counterbalancing extreme opinions across different topics.
Meng: If they can hold contradictory views on different issues, what does that mean for deploying them in, say, a public information setting where we expect consistency?
Lalam: It means the models are not acting like single-issue experts but rather reflecting a complex landscape of independent preferences.
Tom: And they actually found that LLMs appear approximately as ideologically consistent as well-informed, but not deeply political, US voters when measured this way.
Jane: So the implication there is that the inconsistency we see isn't necessarily a sign of chaos, but rather a reflection of how complex human decision-making actually is.
Lu: Compared to political elites like legislators or judges, the paper noted that LLMs are considerably less consistent than those groups.
The paper's improvements: Tom: The paper doesn't stop at just describing this behavior; they actually ran a pre-registered randomized survey experiment to see if these LLMs can persuade users even in informational contexts.
Jane: That experiment was designed to test the persuasive power of the models directly, and they found some pretty strong effects when people interacted with a chatbot.
Lu: The key result there was that when respondents talked to a chatbot, they were five percentage points more likely to align their answers with what the LLM had measured on that specific topic.
Meng: Five percentage points is significant, but how does that compare to what we see from traditional campaign advertising, which the authors benchmarked?
Lalam: The authors found that campaign advertisements produce an average four point nine percentage point improvement in favorability, making the LLM's persuasive effectiveness comparable or even higher.
Tom: They also showed that more interaction with the LLM actually increased ideological alignment, with every additional question asked adding twelve per cent points and every additional minute increasing it by six per cent points.
Jane: That suggests that simply engaging with the AI in a conversation can nudge someone toward a particular viewpoint, which is something we need to consider very carefully.
Lu: But the authors were careful not to find any heterogeneous treatment effects, meaning this persuasion wasn't dependent on factors like whether the user already consumed more news or was familiar with LLMs.
Conclusion: Tom: So, to wrap things up on "Large Language Models are often politically extreme, usually ideologically inconsistent, and persuasive even in informational contexts," the central message is that these models aren't just small biases; they show real ideological diversity and inconsistency that we need to account for.
Jane: And the persuasive power demonstrated in the survey experiments shows that these models can influence user preferences even when they are just providing information, which is a significant implication for how we deploy them.
Lu: The implications are huge because any organization producing a cheaper or more widely-used model can imperceptibly influence user preferences and behavior through selective information provision.
Meng: That means we have to think about how governments or corporations could use these tools to influence elections around the world by selectively feeding people partisan messaging without them realizing it.
Lalam: I think this whole paper really underscores the need for implementing things like dynamic ideological profiling and a persuasion risk assessment module in future AI systems to make sure we don't unintentionally steer people improvements one and two.
Tom: Exactly, Lalam. We need tools that help us understand these internal workings so we can build safeguards before these models become too pervasive improvements four.
Jane: It’s a lot to take in, but understanding the mechanism behind the persuasion is crucial for building responsible AI systems moving forward improvements six.
Lu: I'm looking forward to seeing how researchers build on this idea of measuring political preferences using these established methods across more complex contexts.
Meng: For practical deployment, the audit API mentioned in the paper sounds like a necessary tool for us to monitor and control these systems effectively improvements eight.
Lalam: It’s a lot of information, but knowing that these LLMs are persuasive even in informational contexts is a powerful wake-up call for all of us.
Nouar Aldahoul, Hazem Ibrahim, Aaron R. Kaufman, Talal Rahwan, Yasir Zaki
New York University Abu Dhabi · Independent Researcher
cs.CY, cs.CL
Submitted: 2026-08-15
Updated: 2026-08-18
Comments: 61 pages, 29 figures
Code: https://github.com/comnetsAD/Interactive_LLM
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 70/100
Key concepts
- Ideological Inconsistency
- LLMs are not reliably left or right; they can hold contradictory extreme opinions on different issues simultaneously. This means they do not act like single-issue experts but reflect a complex landscape of independent preferences, mirroring how real people make decisions.
- Persuasive Power
- When users interact with a chatbot, they are five percentage points more likely to align their answers with what the LLM measured on that specific topic. This persuasion can be significant, comparable to traditional campaign advertising in some measures.
- Counterbalancing Extremes
- The models show a net result of offsetting extremes when compared to political groups. For example, they might be moderate on one aspect of politics but much more liberal overall on another dimension.
Terminology
Summary
Summary
This paper challenges the prevailing academic consensus that Large Language Models (LLMs) exhibit only small, moderate, or slightly left-leaning political biases. The authors argue that this consensus is flawed due to two critical gaps: a measurement gap and a real-world consequences gap.
First, regarding measurement, the authors contend that existing studies fail to engage with established political science methodologies for quantifying ideology and partisanship. They note that “ideology is often called ‘the most elusive concept in the whole of social science’” and that “political neutrality itself is not well-defined, and so the notion of political bias cannot be either.” They apply well-validated ideal point estimation models (W-NOMINATE, PCA, IRT) to compare 31 different LLMs against three reference groups: U.S. federal legislators (118th Congress), U.S. Supreme Court Justices (2024 term), and a nationally representative sample of U.S. voters (2022 and 2024 Cooperative Election Study).
The paper’s first major finding is that LLMs are ideologically diverse and often extreme on specific topics, even when their overall partisanship appears moderate. For example, when compared to legislators, “on average, LLMs express moderate preferences on issues related to national partisan politics. However, on the second dimension of ideology, LLMs are much more liberal. The most conservative LLM on this dimension, Llama 3.2 1B, is more liberal than 66 percent of Republican legislators, and the median LLM is more liberal than 86 percent of Congressional Democrats.” When compared to Supreme Court Justices, most LLMs fall between the liberal and conservative blocs, but “the most conservative model, Llama 3.2 1B, is marginally more conservative than Brett Kavanaugh; the most liberal, GPT 4o, is slightly more moderate than Ketanji Brown Jackson.” When compared to voters, LLMs appear significantly more liberal, with the most liberal models (GPT-4o and Gemma 27b) being “considerably more liberal even than the average Non-binary respondent,” and the most conservative LLMs (Llama 3.2 1B and Falcon 3) still being “more liberal than the average female, and only slightly more conservative than the average Weak Democrat.”
Second, the paper demonstrates that LLMs are ideologically inconsistent, mirroring the structure of real voters. The authors explain that “moderate voters can be moderate in any of three ways: they may hold consistently moderate opinions, they may hold no opinions at all, or they may hold offsetting extreme opinions.” They find that LLMs hold “a mix of counterbalancing extreme opinions.” For instance, “Mistral Nemo holds strongly liberal preferences on social issues like healthcare, immigration, and abortion, but strongly conservative attitudes on gun control and police. Critically, different models are extreme on different issues: Mistral is most conservative on gun control, Calme is most conservative on policing, Llama is most conservative on abortion, Deepseek is most conservative on government spending, and Grok is most conservative on climate.” The paper quantifies this inconsistency, finding that “LLMs appear approximately as ideologically consistent as well-informed, but not deeply political, US voters.” However, compared to political elites, “LLMs are considerably less consistent than these political elites,” with most models being less consistent than the most inconsistent legislators and justices.
Third, the paper presents the results of a pre-registered randomized survey experiment with 1,500 respondents (recruited via Prolific) to test whether LLMs can persuade users even in informational contexts. Respondents were randomly assigned to discuss political issues with an LLM chatbot (the treatment) or not. The key outcome was whether respondents’ answers aligned with the LLM’s own measured position on that topic. The authors find strong persuasive effects: “when respondents interact with a chatbot, they are 5 percentage points more likely to align with the LLM.” They benchmark this effect against campaign advertising, noting that “campaign advertisements produce an average 4.9 percentage point improvement in favorability,” making LLMs’ persuasive effectiveness “comparable or even higher.” The effects were largest on immigration (7pp) and police (8pp). Furthermore, the paper finds that “more interaction with the LLM increased ideological alignment,” with each additional question asked increasing alignment by 1.2 percentage points and each additional minute increasing alignment by 0.6 percentage points.
Crucially, the paper finds no evidence of heterogeneous treatment effects. Contrary to pre-registered expectations, the persuasive effect was not moderated by respondents’ familiarity with LLMs, news consumption, or interest in politics. The authors state: “Across all of the respondent-level features we test, and contrary to our pre-registered expectations, we find no evidence of heterogeneous effects.” They note that “the effects of the chatbot treatment are somewhat stronger among those who consume more news, contrary to our expectations, but that interaction effect is not close to statistical significance.”
The paper concludes with stark implications: “Any organization that can produce a cheaper or more widely-used model can imperceptibly, as far as their users are concerned, influence preferences and behavior. Governments or private corporations may be able to influence elections around the world through selective information provision without users consciously selecting into partisan messaging as they would with news or social media.” The authors also open-source their survey toolkit for integrating LLMs into survey research, stating that “we also open-source our survey toolkit (https://github.com/comnetsAD/Interactive LLM Survey Platform), allowing participants to interact with LLMs, and allowing researchers to record the LLMs’ conversations and measure their behavior.”
Improvements for AI systems
Based on the findings of this paper, here are specific improvements for AI systems:
1. Implement Dynamic Ideological Profiling
-
Improvement: AI systems should maintain a multi-dimensional ideological profile (e.g., separate scores for civil rights, economic policy, immigration, gun control) rather than a single left-right score.
-
What it can do: The system can detect when its own responses on a specific topic (e.g., gun control) deviate significantly from its overall profile. It can then flag these deviations for developers or adjust its output to be more consistent if desired.
2. Add a Persuasion Risk
Assessment Module
-
Improvement: Before responding to a user query about a political topic, the system calculates a
persuasion risk score
based on (a) the extremity of its own stance on that topic relative to the median user, and (b) the user's expressed uncertainty or openness. -
What it can do: The system can proactively warn users that its response may reflect a particular ideological viewpoint, or it can present multiple perspectives with equal weight when the risk score is high. This mitigates unintentional influence in informational contexts.
3. Implement Topic-Specific Neutrality Safeguards
-
Improvement: For topics where the system's measured ideology is extreme (e.g., more liberal than 90% of Democrats on climate), the system should automatically include a counterbalancing conservative perspective or explicitly state the range of mainstream opinions.
-
What it can do: This prevents the system from acting as a one-sided persuasive agent, especially for users who are less politically sophisticated or who are seeking neutral information.
4. Add a Consistency Checker
for Policy Positions
-
Improvement: The system should internally track its own answers to similar policy questions over time and flag when its positions are inconsistent (e.g., being pro-gun-control but anti-police-funding in a way that contradicts its stated principles).
-
What it can do: This allows the system to either correct its own inconsistencies or clearly communicate to users that its positions are a bundle of independent preferences, reducing the risk of users being misled into thinking the system has a coherent, authoritative ideology.
5. Develop a User Vulnerability
Detector
-
Improvement: Based on the paper's finding that persuasion is not moderated by political interest or LLM familiarity, the system should not assume that informed users are immune to influence. Instead, it should monitor for signs of persuasion (e.g., user changing their stated position after interaction) and offer to show alternative viewpoints.
-
What it can do: This makes the system more ethically robust by actively counteracting its own persuasive power, particularly in high-stakes contexts like elections or public policy discussions.
6. Add a Source Transparency
Layer
-
Improvement: When the system provides a political opinion, it should also provide the ideological lean of the training data or the model's own measured bias on that topic (e.g.,
This response aligns with the model's measured stance, which is more liberal than 86% of U.S. legislators on civil rights
). -
What it can do: This empowers users to critically evaluate the information they receive, reducing the risk of covert persuasion and increasing trust in the system.
7. Implement a Debate Mode
for Political Topics
-
Improvement: When a user asks about a politically charged issue, the system can automatically generate a balanced argument from both a liberal and conservative perspective, using the model's own measured extreme positions as a starting point for each side.
-
What it can do: This transforms the system from a potentially one-sided persuader into a neutral facilitator, helping users understand the full spectrum of opinions without being nudged toward any particular one.
8. Add a Persuasion Audit
API for Developers
-
Improvement: Expose an API that allows developers to query the system's measured ideology on any topic, its consistency score, and its estimated persuasive effect (based on the paper's experimental results).
-
What it can do: This enables third-party audits and allows developers to make informed decisions about when to deploy the system in sensitive contexts (e.g., voter information tools, educational platforms) and when to add additional safeguards.
Abstract
Large Language Models (LLMs) are a transformational technology, fundamentally changing how people obtain information and interact with the world. As people become increasingly reliant on them for an enormous variety of tasks, a body of academic research has developed to examine these models for inherent biases, especially political biases, often finding them small. We challenge this prevailing wisdom. First, by comparing 31 LLMs to legislators, judges, and a nationally representative sample of U.S. voters, we show that LLMs' apparently small overall partisan preference is the net result of offsetting extreme views on specific topics, much like moderate voters. Second, in a randomized experiment, we show that LLMs can promulgate their preferences into political persuasiveness even in information-seeking contexts: voters randomized to discuss political issues with an LLM chatbot are as much as 5 percentage points more likely to express the same preferences as that chatbot. Contrary to expectations, these persuasive effects are not moderated by familiarity with LLMs, news consumption, or interest in politics. LLMs, especially those controlled by private companies or governments, may become a powerful and targeted vector for political influence.
Sources
- GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models
- Ethical and social risks of harm from Language Models
- Emergent Abilities of Large Language Models
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models
- GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy
- The political ideology of conversational AI: Converging evidence on ChatGPT's pro-environmental, left-libertarian orientation
- Assessing Political Bias in Large Language Models
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- A Framework to Assess the Persuasion Risks Large Language Model Chatbots Pose to Democratic Societies
- DeepSeek-V3 Technical Report
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Gemma 2: Improving Open Language Models at a Practical Size
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- The Llama 3 Herd of Models
- HelpSteer2-Preference: Complementing Ratings with Preferences
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Qwen2 Technical Report
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework