Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases
McGill University · University Canada West · Toronto Metropolitan University
cs.CL
Submitted: 2026-08-11
Updated: 2026-09-07
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: This paper introduces an analytically exact framework for the controlled behavioral evaluation of Large Language Models (LLMs), shifting from large, unstructured benchmarks to fully crossed factorial
Terminology
Summary
This paper introduces an analytically exact framework for the controlled behavioral evaluation of Large Language Models (LLMs), shifting from large, unstructured benchmarks to fully crossed factorial experiments. The framework addresses three methodological gaps: experimental design, measurement, and analysis.
Design Gap: The paper replaces unstructured prompting with fully crossed factorial experiments to systematically isolate causal main and interaction effects. The authors state: "We formulate LLM evaluation as a fully crossed factorial design, where every distinct variable (e.g., demographic, context, LLM persona) is treated as an independent factor. We systematically evaluate every possible combination using validated psychometric scales."
Measurement Gap: The framework eliminates Monte Carlo text sampling noise by operating directly on exact, token-level Probability Mass Functions (PMFs). The paper explains: "An LLM possesses a constant, fully observable internal state for a given prompt. Generating multiple text responses is effectively taking repeated, noisy draws from a static mathematical distribution. Our framework, instead, operates directly on exact, token-level Probability Mass Functions (PMFs) to eliminate sampling error."
Analysis Gap: The paper derives a multivariate ordinal consensus metric and a distributional ANOVA to process these PMFs analytically. The authors note: "We bridge this gap by introducing a multivariate consensus metric that respects ordinal distance. We also formulate distributional ANOVA to analytically compute the exact probability distributions of the isolated causal effects."
The methodological core includes:
-
A distributional Hoeffding decomposition with Optimal-Transport-based paired contrasts and a formal ANOVA-consistency guarantee (Theorem 3.1), which
decomposes the composite construct distribution into baseline, main-effect, and interaction-effect distributions.
-
Propagation of item-level PMFs into an exact composite-score distribution using discrete convolution.
-
A multivariate extension of the Consensus metric (Tastle and Wierman, 2007) to address the failure mode of Shannon entropy on Likert scales, since
entropy cannot distinguish between an LLM heavily polarized between 'Strongly Agree' (5) and 'Strongly Disagree' (1) versus one smoothly concentrated around 'Neutral' (3).
The framework is validated with a case study on consumer ethnocentrism using the 17-item CETSCALE across five LLMs (Llama 3.3 70B, Gemma 3 27B, Qwen3 Next 80B, Aya Expanse 32B, Ministral 14B) and four target countries (USA, China, Canada, France), yielding 340 exact item-level predictive distributions.
Key findings from the case study:
-
Constraint: Most models show near-perfect adherence (failure rate < 0.001), but Aya Expanse shows the highest aggregate failure rate (M = 0.006) with large variance.
-
Consensus: Four models show high internal agreement (Cns > 0.88), but Ministral is a major outlier (Cns ≈ 0.65–0.67), showing severe behavioral polarization.
-
Main Effects: Aya-32B exhibits a severe positive deviation from the grand mean (E = 21.19, SNR = 1.63, dPD > 0.99), while Qwen3-80B demonstrates a robust negative main effect (E = −14.63, SNR = 1.20, dPD = 0.93). For target countries, there is a systemic decrease in ethnocentrism when the target is China (E = −6.46, SNR = 1.04).
-
Interaction Effects: The paper defines country-of-origin bias as a positive interaction between a model and its own country. US-developed models (Gemma 3 and Llama 3.3) show positive own-country interactions (+3.21 and +2.59), but this is not universal: Aya-32B (Canada, −3.77) and Qwen3 (China, −0.20) do not exhibit positive own-country interactions.
The paper also demonstrates the cost of generative sampling: "At N = 10, standard sampling flips the sign of small effects (E < 1) 18% of the time, and even at N = 100, it flips them 6.3% of the time." Additionally, non-default decoding (temperature scaling) can artificially shift latent construct scores by up to 2.79 points for low-consensus models.
Finally, the framework is extended to prompt sensitivity by promoting the administration prompt to a first-class experimental factor, showing that prompt framing acts as a causal main effect of its own, shifting the latent construct score by up to 13.5 points depending on the phrasing used,
while the core country-of-origin signals remain robust under statistical control for framing.
Improvements for AI systems
Improvements to AI systems:
-
Exact-output evaluation mode: Implement a
distributional inference
API that returns the exact token-level Probability Mass Function (PMF) for any prompt, instead of sampling text. This eliminates Monte Carlo noise, allowing downstream systems to compute stable, reproducible scores for any behavioral metric without repeated generation. -
Causal factor isolation for prompt engineering: Use the fully crossed factorial design to automatically identify which prompt components (e.g., demographic framing, persona, instruction wording) causally drive output shifts. The improved system can report main and interaction effects for any prompt variable, enabling precise, interpretable tuning rather than trial-and-error.
-
Ordinal-aware consensus monitoring: Replace entropy-based uncertainty metrics with the multivariate ordinal consensus metric. The improved system can detect pathological polarization (e.g., an LLM oscillating between extreme agree/disagree) versus genuine neutrality, flagging models that are unreliable for opinion-sensitive tasks.
-
Decoding-parameter sensitivity alerts: Before deployment, the system can compute how much non-default decoding (e.g., temperature scaling) shifts latent construct scores. It will warn users when a model's low consensus makes it vulnerable to artificial score shifts (up to 2.79 points), recommending default decoding for stable outputs.
-
Prompt-framing bias correction: The system can treat the administration prompt as a first-class experimental factor, estimate its causal main effect on any measured construct, and statistically control for it. This yields bias-corrected scores that remain robust across phrasings, improving fairness in comparative model evaluations.
-
Small-effect sign-flip prevention: For any effect size estimate, the system can compute the minimum number of samples needed to avoid sign flips (e.g., >18% flip rate at N=10 for E<1). It will automatically request more samples or switch to exact PMF analysis when effect sizes are small, ensuring reliable directional conclusions.
-
Own-country bias detection: The system can run a targeted interaction analysis to detect whether a model exhibits positive bias toward its developer’s country of origin. This enables auditing for hidden cultural biases in deployed systems, allowing developers to correct or disclose such biases.
What the improved AI system can do:
-
Provide exact, noise-free behavioral measurements for any prompt, without sampling variance.
-
Isolate causal drivers of output (model, context, persona, prompt framing) with statistical guarantees.
-
Detect and correct for ordinal-scale polarization, decoding sensitivity, and prompt-framing artifacts.
-
Produce bias-adjusted scores for cross-model and cross-country comparisons, with confidence intervals derived from exact distributions.
-
Automatically flag unreliable outputs when sample sizes are insufficient or when interaction effects (e.g., own-country bias) are present.
Abstract
As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks. While effective for general capabilities, these datasets fundamentally conflate causal mechanisms: even when an aggregate bias is detected, unstructured evaluations cannot disentangle whether it stems from baseline traits, contextual confounders, or complex interactions. To address this, we introduce an analytically exact framework for the controlled behavioral evaluation of LLMs. We bridge human psychometrics with LLM mechanics by resolving gaps in design, measurement, and analysis. First, we replace unstructured prompting with fully crossed factorial experiments to systematically isolate causal main and interaction effects. Second, we eliminate Monte Carlo text sampling noise by operating directly on exact, token-level Probability Mass Functions (PMFs). Third, we derive a multivariate ordinal consensus metric and a distributional ANOVA to process these PMFs analytically. We validate our framework with a case study on consumer ethnocentrism across five LLMs, demonstrating how our approach isolates systemic country-of-origin biases that aggregate benchmarks otherwise obscure.
Sources
- Value Drifts: Tracing Value Alignment During LLM Post-Training
- Measuring Massive Multitask Language Understanding
- Language Models (Mostly) Know What They Know
- Is It Bad to Work All the Time? Cross-Cultural Evaluation of Social Norm Biases in GPT-4
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering