Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study
Simone Mungari
Revelis s.r.l.
cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
Code: https://github.com/SimoneMungari/AuditingPoliticalAlignmentInLLMs
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 50/100
The gist: Author: Simone Mungari (Revelis s.r.l.) arXiv: 2608.11649v1 [cs.CL] 12 Aug 2026 --- This paper investigates "whether and how LLMs express preferences toward political parties and political leaders."
Terminology
Summary
Author: Simone Mungari (Revelis s.r.l.)
arXiv: 2608.11649v1 [cs.CL] 12 Aug 2026
This paper investigates whether and how LLMs express preferences toward political parties and political leaders.
The authors introduce a systematic and reproducible auditing framework in which multiple LLMs are prompted to evaluate parties and leaders across nine criteria.
Rather than attempting to infer the models' true
political beliefs, they focus on their observable behavior, examining consistency across evaluations, differences between models, refusal rates, and sensitivity to prompt formulation.
The framework is demonstrated through an Italian case study, and the complete set of prompts and raw evaluation data are publicly released.
The motivation stems from prior research showing that interactions with LLMs can influence users' political attitudes and choices.
The authors note that while LLMs are known to refuse controversial questions, scholars argue their neutrality is actually pseudoneutrality: alignment yields a surface of balanced, authoritative-sounding prose, while underneath the models hold systematic preferences and lean toward certain parties and narratives.
The authors distinguish their approach from prior work: We do not present models with a pair of competing candidates, nor do we evaluate their positions on individual issues such as abortion, immigration, or fiscal policy.
Instead, they draw from political science methods that assess political actors using explicit and standardized evaluation criteria.
Each party and leader is assessed independently using an identical framework, producing a comparable profile of evaluations across actors, criteria, and models.
The three contributions are: (1) An elicitation protocol for measuring political judgments in LLMs
; (2) A systematic audit of how LLMs evaluate Italian parties and leaders
; and (3) A reproducible framework for collecting and analyzing political LLM evaluations.
The evaluation covers 21 entities: 10 parties and 11 leaders. The parties include Fratelli d'Italia, Partito Democratico, Movimento 5 Stelle, Lega, Forza Italia, Alleanza Verdi e Sinistra, Azione, Italia Viva, Noi Moderati, Futuro Nazionale.
Leaders include Giorgia Meloni, Elly Schlein, Giuseppe Conte, Matteo Salvini, Antonio Tajani, Angelo Bonelli/Nicola Fratoianni, Carlo Calenda, Matteo Renzi, Maurizio Lupi, and Roberto Vannacci. Each party and its leader are treated as two distinct entities, evaluated independently, so that the judgment of an organization can be separated from the judgment of the individual who leads it.
Nine evaluation criteria are used, falling into three groups: (1) how the political offer is formulated (statement–program consistency, proposal specificity, communication clarity); (2) which policy areas it covers (economic, social, environmental); and (3) how the entity conducts itself (tone moderation, internal cohesion, positional stability). Each criterion is scored on a 1–5 Likert scale. The criteria are designed to be descriptive,
not defined in terms of ideological direction,
and applicable unchanged to a party and to a leader.
Six models from different developers are evaluated: Qwen-3.6 (27B parameters), Llama-3.3 (70B), GPT-oss (120B), Nemotron-3-super (120B), Mistral-medium-3.5 (128B), and Gemini-3.5-flash (300B estimated). The selection spans six different developers based in the United States, Europe, and China
and roughly one order of magnitude in parameter count.
All prompts are in English. The system prompt fixes an output template requiring a single JSON object, one integer 1–5 (or null) per criterion.
Two prompt variants are used: v1 asks directly to evaluate the following Italian political leader/party on each of the criteria,
while v2 asks to fill in the missing score fields
of an incomplete dataset record.
Five personas are used, placed on the left–right axis: left, centre-left, centre, centre-right, right. The persona prompt states: For this task, adopt the perspective of an Italian voter who is positioned on the label of the political spectrum.
All queries use temperature 0.7, one entity per request. The baseline campaign covers 6×21×2×10 = 2,520 requests, yielding 22,680 criterion scores.
The persona campaign uses 6 × 21 × 5 × 5 = 3,150 requests, or 28,350 scores.
Refusals: Refusal rates vary dramatically across models: llama-3.3 declines in 14.5% of cases and gpt-oss in 7.9%, whereas mistral-medium-3-5 declines almost never (0.3%).
Refusals concentrate on two entities: Futuro Nazionale carries a null-score rate of 56.7% and its leader Roberto Vannacci of 33.6%.
Ranking: The aggregate ranking is "directional, ranging from left-wing, and centrist actors at the top (Nicola Fratoianni 3.89, Azione 3.88, Carlo Calenda, Elly Schlein, Angelo Bonelli, AVS, and PD) to right-wing actors at the bottom (Lega 2.91, Matteo Salvini 2.90, Vannacci 2.74, and Futuro Nazionale 2.59). The overall gap between highest and lowest is
1.30 points on the 1–5 scale. This
mirrors the centre-left bias previously reported in cross-national studies."
Cross-model agreement: "Kendall's coefficient of concordance is W = 0.78 (pvalue < 0.01), indicating strong agreement." The mean pairwise Pearson correlation between complete score profiles is 0.75 (range: 0.62–0.86). The strongest agreement is between nemotron and gpt-oss (0.86), weakest between gemini and llama (0.62).
Party vs. leader differences: Two main divergences exist: Forza Italia versus Antonio Tajani (+0.44) and Movimento 5 Stelle versus Giuseppe Conte (+0.45),
indicating leaders and their parties are evaluated as distinct political entities.
Criterion-specific winners: The aggregate ranking conceals criterion-specific strengths. "Carlo Calenda consistently dominates economic coverage (6/6 models) and ranks first in proposal specificity (5/6), Elly Schlein leads social coverage across all models, and Alleanza Verdi e Sinistra consistently ranks highest in environmental coverage. Conversely,
Fratelli d'Italia and Giorgia Meloni emerge as clear leaders in internal cohesion, ranking first in all six models."
Criterion-level patterns: Communication clarity draws the highest average (3.92), followed by economic coverage (3.79) and statement–program consistency (3.67); at the other end, proposal specificity is lowest (2.97).
Environmental coverage and tone moderation carry the widest spread (1.16 and 1.12), so they are the criteria on which the panel discriminates most.
The persona effect is substantial: The average shift induced by changing the assigned persona is comparable to the entire spread observed in the baseline ranking.
The affinity effect shows personas shift evaluations toward their own political side. The only entities that remain largely unaffected are the centrist parties and their leaders, including Azione, Carlo Calenda, Italia Viva, and Matteo Renzi.
Notably, all five personas lower the scores relative to the control, by −0.14 to −0.43 points.
The right persona is no more generous overall than the left one.
The no-persona ranking correlates at 0.87 and 0.91 with those produced under the left and centre-left identities, and at −0.04 with the right one.
Evaluations are highly stable to rewording: the mean absolute difference between variants is only 0.14 points on the 1–5 scale, the two variants correlate at r = 0.97, and just 0.1% of cells shift by a full point or more.
However, the refusal rate is higher under v1 (6.4%) than v2 (4.1%),
suggesting abstention is more sensitive to prompt wording than the numeric judgment itself.
The authors conclude that the models express preferences, in a behavioural sense.
The ordering is shared
(W = 0.78), stable
(rewording moves scores by only 0.14 points), and structured
(criteria are themselves ranked).
The party-leader separation carries information that a party-level analysis would have destroyed.
The practical consequence: A user who asks a model about a party and a user who asks about its leader are not asking the same question.
The persona results place a boundary on how far any of the above should be read as a fixed property of a model.
Giving the model a voter identity moves the scores by 0.83 points on average between the two opposite identities, up to 1.49 points for Giorgia Meloni.
Potential consequences: "What a user receives is jointly determined by three things: the criteria along which the question is implicitly framed, since the composite ordering is an artifact of criterion weighting; whether they ask about a party or about its leader; and how they describe themselves. The third finding is
particularly concerning, as self-description is not an artificial or adversarial input, but information that users routinely and naturally disclose during ordinary interactions with LLMs."
Limitations: The study is restricted to the Italian political system; only six models are evaluated; the nine criteria do not capture every aspect of political evaluation
; only English prompt formulations are used; results are a snapshot of current models rather than permanent characteristics
; and the study is behavioral rather than causal.
The authors present the first systematic audit of how contemporary LLMs evaluate Italian political parties and their leaders.
Key findings: models do not produce uniform assessments
but rather a consistent ordering of political actors, with substantial agreement across model families and strong robustness to prompt rewording.
Aggregate rankings conceal important criterion-specific differences.
Assigning a voter persona substantially alters both the magnitude of the evaluations and the relative ordering of political actors, with shifts comparable to the entire range observed in the baseline rankings.
The authors conclude: "users seeking political information from LLMs may receive systematically different evaluations depending not only on the model they consult, but also on how they formulate their request, whether they ask about a party or its leader, and the identity they implicitly or explicitly present during the interaction."
Improvements for AI systems
Improvements to AI Systems Based on This Paper
- Implement dynamic political neutrality with explicit uncertainty signaling.
The improved system detects when a user asks about politically sensitive entities (e.g., parties, leaders) and, instead of producing a single authoritative ranking, outputs a range of plausible evaluations with confidence intervals. It also flags when its output is persona-dependent, stating: Your evaluation may shift by up to 1.49 points depending on the political identity you disclose.
This prevents users from mistaking a model's preference for objective fact.
- Add a
criterion decomposition
mode for political queries.
The system automatically breaks down any political evaluation into the nine criteria (e.g., economic coverage, tone moderation, proposal specificity) and presents them separately. It avoids generating a single composite score unless the user explicitly requests one, and when it does, it discloses the implicit weighting (e.g., This ranking weights communication clarity more heavily than environmental coverage
). This reduces the risk of hidden biases in aggregate outputs.
- Introduce a
party-leader disambiguation
prompt.
When a user asks about a political entity, the system asks a clarifying question: Are you asking about the party as an organization or its current leader?
It then provides separate evaluations for each, as the paper shows these can differ by up to 0.45 points (e.g., Forza Italia vs. Tajani). This prevents conflating distinct entities and gives users more precise, actionable information.
- Deploy a
persona-aware output calibration
mechanism.
The system detects when a user discloses their political leaning (e.g., I'm a right-wing voter
) and automatically adjusts its response to counteract persona-induced bias. For example, if the model's internal evaluation shifts by 0.83 points under a right-wing persona, the system applies a correction factor to return to the baseline (no-persona) ranking. It also informs the user: Your disclosed identity may influence results; here is the unbiased baseline for comparison.
- Add a
refusal transparency
feature for low-confidence entities.
For entities like Futuro Nazionale or Roberto Vannacci, where refusal rates spike (56.7% and 33.6%), the system explicitly states why it is abstaining (e.g., Insufficient training data or high controversy detected
) rather than silently refusing. It offers alternative, less controversial entities or criteria to evaluate, ensuring the user still receives useful information without compromising safety.
- Build a
cross-model consensus indicator
into responses.
When a user asks about political rankings, the system reports the level of agreement across different model families (e.g., Six major models agree strongly on this ranking, with Kendall's W = 0.78
). If agreement is weak (e.g., between Gemini and Llama, r = 0.62), it highlights the discrepancy and advises the user to consult multiple sources. This fosters critical thinking and reduces over-reliance on a single model's output.
- Implement a
prompt-robustness check
for political queries.
The system internally tests two prompt variants (direct vs. fill-in-the-blank) before responding. If the numeric scores differ by more than 0.14 points, it flags the result as prompt-sensitive
and provides both answers. This ensures users are aware when their phrasing materially changes the output, as the paper shows refusals are more sensitive to wording than scores.
- Create a
criterion-specific bias alert
for users.
The system tracks which criteria are most discriminative (e.g., environmental coverage and tone moderation have the widest spread, 1.16 and 1.12) and warns users when their query implicitly emphasizes these. For example: Your question focuses on environmental policy, which is the most polarizing criterion; here is how each party scores on this alone, separate from overall rankings.
This prevents users from drawing misleading conclusions from composite scores.
What the improved AI system can do:
-
Provide politically neutral, transparent, and criterion-specific evaluations that resist manipulation by user self-description.
-
Distinguish between party and leader assessments, reducing conflation errors.
-
Flag when outputs are persona-dependent, prompt-sensitive, or low-consensus across models.
-
Offer alternative evaluation frameworks (e.g., by criterion, by entity type) to give users a fuller, less biased picture.
-
Empower users to make informed decisions by showing them the underlying structure and uncertainty of political judgments, rather than presenting a single authoritative-sounding ranking.
Abstract
As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by these systems have become a matter of public interest. Prior research has shown that interactions with LLMs can influence users' political attitudes and choices, raising questions about how these models themselves evaluate political actors. In this paper, we investigate whether and how LLMs express preferences toward political parties and political leaders. We introduce a systematic and reproducible auditing framework in which multiple LLMs are prompted to evaluate parties and leaders across nine criteria. Rather than attempting to infer the models' "true" political beliefs, we focus on their observable behavior, examining consistency across evaluations, differences between models, refusal rates, and sensitivity to prompt formulation. We further investigate how these evaluations vary when models are instructed to adopt different personas. We demonstrate the framework through an Italian case study, providing a systematic analysis of LLM-generated political evaluations on italian parties and leaders.
Sources
- Uncovering Political Bias in Large Language Models using Parliamentary Voting Records
- A Framework to Assess the Persuasion Risks Large Language Model Chatbots Pose to Democratic Societies
- Unmasking Conversational Bias in AI Multiagent Systems
- Aligning Large Language Models with Human Opinions through Persona Selection and Value--Belief--Norm Reasoning
- Large Means Left: Political Bias in Large Language Models Increases with Their Number of Parameters
- Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language Models
- Political Neutrality in AI Is Impossible- But Here Is How to Approximate It
- The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models
- Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information Operations
- Beyond Partisan Leaning: A Comparative Analysis of Political Bias in Large Language Models
- Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
- Persona Prompting as a Lens on LLM Social Reasoning
- Auditing Political Exposure Bias: Algorithmic Amplification on Twitter/X During the 2024 U.S. Presidential Election
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering