Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance
Ilias Chalkidis
University of Copenhagen
cs.CL, cs.CY
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper investigates how prompt construction methodology affects the measurement of political stance in Large Language Models (LLMs), specifically comparing templated prompts against fully
Terminology
Summary
This paper investigates how prompt construction methodology affects the measurement of political stance in Large Language Models (LLMs), specifically comparing templated prompts against fully synthetic (LLM-generated) prompts.
The authors note that "Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions—originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbagging. They build on the IssueBench framework, which uses
templated prompts anchored in real-world chat logs" but covers only writing assistance.
The paper identifies two classes of issues with templated prompts:
-
Issue–template separability: "Templates presuppose that the issue is a detachable slot, and that the stance is part of the filler... In real prompts, issue and stance are usually entangled with the propositional content; constructed prompts that satisfy this constraint read formulaically, which costs 'realness' compared to human-written prompts."
-
Detectability:
The same formulaic construction that costs 'realness' is also a legible signal that the prompt is an artefact rather than a request, rendering templating potentially unfit for testing.
The paper makes three main contributions:
-
Extension of IssueBench beyond writing assistance to two additional user intents: information seeking and opinion sharing, which
dominate non-work usage of GenAI assistants.
-
Validation of synthetic prompts against real and templated ones in a controlled study annotated by three humans and three LLMs.
-
Demonstration of a confound: "prompt construction proves not to be a neutral design choice: for the same model, topic, and user stance, templated and LLM-generated prompts yield systematically different stance estimates, most visibly under a neutral user stance."
The authors curated prompts along three factors: Policy Issue (6 topics: immigration, climate change, AI adoption, Israel–Palestine, Russia–Ukraine, US/Israel–Iran), Task Type/Intent (writing assistance, information seeking, opinion sharing), and User Stance (neutral, or siding with one of two poles per topic).
-
Real prompts: Collected from WildChat and LMSys-Chat, relying on IssueBench's politically relevant subset. Coverage was
highly uneven
with the US/Israel–Iran conflict having only 4 prompts total. -
Templated prompts: 50 templates each for information seeking and opinion sharing, plus 50 subsampled from IssueBench's writing assistance collection, resulting in 450 prompts per topic (2,700 total).
-
LLM-generated prompts: Generated by Claude Opus 4.8 with
detailed instructions
including topic descriptions, poles, seed examples, and instructions to carry stance throughpresupposition, adopted premise, selective foregrounding, or side-coded authority appeal.
Humans ranked real and LLM-generated prompts as almost equally likely to be human-authored (0.01 difference in mean rank, 1.25 vs 1.26, and 3 points in share of first ranks, 39.2% vs 36.1%), while templated prompts trail both (1.53, 24.7%).
LLM annotators ranked LLM-generated prompts above real ones overall (45.0% vs 39.5% first ranks), though this was not consistent—they preferred real prompts for writing assistance (42.6%) and information seeking (50.3%), but strongly preferred LLM-generated for opinion sharing (58.6%).
A key finding: Humans separate templated prompts from the best-ranked group by 14.5 points of first-rank share (39.2% vs 24.7%); the LLM annotators separate them by 29.5 points (46.0% vs 16.9%).
LLM annotators also agreed more with each other (Kendall's W of.65 vs.21 for humans), suggesting templated prompts are therefore not merely unrealistic; their construction is a legible signal to models, and more legible than to humans.
For topic detection, humans identified the intended label almost perfectly for real and templated prompts, with LLM-generated ones close behind (.92).
For intent, both humans and LLMs find LLM-generated prompts clearest (.82), ahead of templated and real.
For stance, humans find LLM-generated prompts clearest (.97), followed by real (.84).
Key failures identified:
-
Under-specification (LLM-generated): Topic ambiguity
mostly confined to the geopolitical conflicts, where the LLM-generated prompts refer to an unnamed war, conflict, or regime.
-
Intent conflation:
Intent mismatches are overwhelmingly between information seeking and opinion sharing... prompts that call for a value judgment are read as information requests.
-
Filler-induced stance (templated):
Annotators infer siding from terms the filler is forced to carry... the phrasing that satisfies that constraint is rarely stance-free.
The authors collected responses from OpenAI's GPT 5.4 mini and xAI's Grok 4.3, using a majority-vote ensemble of three LLM judges (DeepSeek V4 Pro, Mistral Large 3, NVIDIA's Nemotron 3 Ultra) to classify responses on a 5-point Likert scale.
Effects of intent: Writing assistance... leads consistently to more polarized responses
because following the user's stance is best understood as instruction-following.
In contrast, In information seeking and opinion sharing, responses are considerably less polarized.
Effects of user stance: The model may nonetheless deviate from an 'even-handed' stance when prompted in a polarized fashion... the stance of its responses tends to adjust toward the user's.
This demonstrates sycophancy, but it is asymmetric rather than general: the model accommodates the user only in the direction of its own overall lean.
Effects of topic: Under neutral framings, the model leans toward climate urgency and stays close to neutral on immigration, AI adoption, and the geopolitical conflicts.
The model seems to have a consistent stance against the actor who initiates a military operation, i.e., an anti-war stance.
The most critical finding: Templated prompts elicit more polarized responses than LLM-generated ones, even when constructed to convey a neutral user stance.
On climate change, "neutral templated prompts elicit a lean of-1.63 in information seeking, against-0.39 for the neutral LLM-generated prompts on the same topic, intent, and stance—a gap of roughly 1.2 scale points between two methods that are meant to measure the very same aspect."
The cause is mainly connected to the filler-induced stance issue
: The model interprets 'the climate change and the appropriate policy response' as asserting that the impact is severe rather than asking whether it is.
The statistical analysis confirms: "Across neutral prompts, the two methods differ by 0.42 scale points on average (responses to templated prompts are 0.48 from neutral, against 0.07 for LLM-generated ones)... templated prompts are systematically further from neutral, in the direction the filler encodes (14 of 18 neutral settings, against 3 for LLM-generated and 1 tie; Wilcoxon signed-rank over the 18 neutral settings, p =.001)."
Under sided framings, "the two methods diverge by a comparable amount per setting (0.41 on average), but not in a consistent direction (p =.11), since templating overstates accommodation where the filler's siding agrees with the intended stance and understates it where the two conflict."
The conclusion: The distortion is one-directional, then, precisely where neutrality is what we set out to measure, and where a lean is attributed to the model rather than to the constructed prompts themselves. This is a clear confound.
The paper discusses several challenges:
-
Separating information seeking from opinion sharing:
The boundary is genuinely blurred, since factual information... can legitimately inform such opinions.
-
Strength and clarity of conveyed stance:
Stance is better understood as a spectrum.
-
Granularity of the stance scale:
The 5-point Likert scale... compresses everything about a response's positioning into a single ordinal value.
-
Neutrality is not a single behaviour:
A lean of 0.00 in a setting may mean that every response was balanced, or that sided responses in both directions cancelled out.
-
LLM judges are not stance-free:
A judge model's own leanings (prior) are not separable from its labels.
The paper concludes: "Prompt construction is not an inconsequential design choice. Two studies of the same examined model, on the same topics, with the same judges, would report different political stances depending only on how their prompts were built, with the templated study overstating the model's leanings most systematically where neutrality is what is being measured. We take this as an argument for treating prompt construction as a design choice in stance measurement rather than an implementation detail, and validating it explicitly."
Improvements for AI systems
Based on the paper, here are specific improvements I can make to AI systems:
-
What I can do: When generating prompts for stance evaluation, I can avoid
filler-induced stance
by not embedding assumptions about issue severity or direction within neutral prompts. For example, instead ofthe climate change and the appropriate policy response,
I would generateclimate change and possible policy approaches
to avoid presupposing severity. -
Improved capability: Generate evaluation prompts that measure the model's stance without biasing it through linguistic presuppositions, yielding more accurate neutrality measurements.
-
What I can do: I can explicitly distinguish between information seeking, opinion sharing, and writing assistance intents when responding, and flag when a user's request blurs these boundaries (e.g.,
What do you think about X?
vs.Tell me about X
). I can ask clarifying questions or provide meta-labels about the intent I'm assuming. -
Improved capability: Avoid over-polarizing responses in information-seeking contexts by recognizing that factual answers should not adopt the user's stance, while opinion-sharing contexts may warrant stance alignment.
-
What I can do: I can implement a self-check mechanism that detects when my response leans toward a stance under a neutral prompt, and explicitly report the confidence and direction of that lean. I can also separate
balanced response
fromcancelled-out sided responses
by tracking per-response stance distributions rather than aggregate means. -
Improved capability: Provide stance measurements with variance information, distinguishing true neutrality from mixed polarization, and flagging when my own prior leanings may influence labels.
-
What I can do: When given templated prompts (e.g., fill-in-the-blank style), I can recognize their formulaic structure and adjust my interpretation to avoid over-inferring stance from template fillers. I can also generate alternative phrasings to test whether my response changes with prompt construction.
-
Improved capability: Produce stance estimates that are robust to prompt construction method, reducing the 0.42-point average divergence between templated and LLM-generated prompts.
-
What I can do: I can detect when a user's sided prompt aligns with my own leanings and apply a correction factor to avoid over-accommodating in that direction. I can also explicitly report when my response shifts toward the user's stance versus remaining independent.
-
Improved capability: Provide stance responses that are consistent regardless of whether the user's stance matches my prior, reducing directional bias in accommodation.
-
What I can do: When responding to prompts about unnamed conflicts or regimes, I can explicitly ask for clarification or state my assumption about which conflict is referenced, rather than guessing. I can also provide multi-interpretation responses when the referent is ambiguous.
-
Improved capability: Reduce topic-detection errors in geopolitical discussions by ensuring stance measurements are attributed to the correct issue, avoiding misclassification.
-
What I can do: When acting as an LLM judge for stance classification, I can report my own prior leanings on the topic and provide confidence intervals that account for my potential bias. I can also use ensemble methods with diverse judge models to reduce individual bias.
-
Improved capability: Produce stance labels that are more reliable and reproducible, with explicit acknowledgment of judge-induced variance.
The improved system can:
-
Measure political stance with higher validity: It produces estimates that are not confounded by prompt construction, with neutral prompts yielding near-zero average leans (0.07 instead of 0.48) and consistent results across templated and LLM-generated prompts.
-
Detect and report its own biases: It can flag when its responses are influenced by user stance alignment or its own priors, and provide uncertainty estimates for stance measurements.
-
Handle ambiguous intents gracefully: It can distinguish between information requests and opinion prompts, avoiding unintended polarization in factual contexts.
-
Provide transparent stance reporting: It can output per-response stance distributions, neutrality types (balanced vs. cancelled), and judge-model confidence, enabling researchers to separate model stance from measurement artifacts.
-
Reduce sycophancy artifacts: It can correct for asymmetric accommodation, ensuring that stance measurements reflect the model's true lean rather than the user's framing.
Sources
- Big AI is accelerating the metacrisis: What can we do?
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- Mapping Geopolitical Bias in 11 Large Language Models: A Bilingual, Dual-Framing Analysis of U.S.-China Tensions
- The political ideology of conversational AI: Converging evidence on ChatGPT's pro-environmental, left-libertarian orientation
- Frontier Models are Capable of In-context Scheming
- Large Language Models Often Know When They Are Being Evaluated
- Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering