Measuring Human-Like Bias in LLMs? A Critique of Human-Derived Bias Constructs in LLM Evaluation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Measuring Human-Like Bias in LLMs? A Critique of Human-Derived Bias Constructs in LLM Evaluation".
Tom: Researchers increasingly use human-derived bias constructs to study Large Language Models (LLMs), including social-cognitive constructs such as implicit bias and stereotype activation, and cognitive biases such as anchoring,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: The core idea of this paper, "Measuring Human-Like Bias in LLMs? A Critique of Human-Derived Bias Constructs in LLM Evaluation," is that when we try to use psychological concepts like implicit bias or anchoring to study Large Language Models, we run into a problem because the way those constructs are measured on humans doesn't transfer directly to how an AI system produces results.
Jane: Exactly. The authors identify this as an inferential gap that shows up because human-derived constructs are being mapped onto very different model-side observables, such as token probabilities or generated completions, which aren't always the right evidence for the claim we want to make about the model itself.
Lu: They lay out three main sources for this mismatch: construct mismatch, where psychological constructs don't neatly fit into the learning processes of an AI; subject mismatch because we are comparing a human mind to a trained artifact; and evaluation context mismatch, which deals with how prompts and decoding choices affect what we measure.
Meng: So, if I understand it correctly, the paper is saying that seeing a certain probability score doesn't automatically prove the model will cause real-world harm in a hiring situation; it’s just an association within the model's internal workings, which is a key distinction for me as someone who builds these systems.
Lalam: I think what they are really highlighting is that we need to be much more careful about what kind of claim we can actually support when we look at LLM outputs; it’s not just about finding a correlation, it's about understanding the nature of that correlation.
Tom: That makes sense. The paper is essentially providing an analytical framework to clarify what kinds of claims different evaluation designs can actually support, rather than just assuming one type of evidence proves another.
Conclusion: Jane: Thinking about the title, "Measuring Human-Like Bias in LLMs? A Critique of Human-Derived Bias Constructs in LLM Evaluation," it seems the main point is that we shouldn't just take human psychological tools and slap them onto AI tests without understanding the specific differences between human cognition and model behavior.
Lu: The authors want us to use a framework to clearly distinguish between different types of findings, like distributional association versus psychological analogy, so researchers don't accidentally overstate what the AI is actually doing.
Meng: For practical implications, this means that instead of just looking at one metric for bias in an LLM evaluation, we need a checklist to determine if that finding points toward simple statistical correlation or if it genuinely suggests a deeper problem with how the model is operating in a way that mirrors human social cognition.
Lalam: If we adopt this framework, it allows us to move beyond just noticing patterns and start making more targeted improvements to how we fine-tune and align these models for better fairness.
Tom: It really boils down to being very precise about what we claim when we present results; the paper's contribution is forcing a much more thoughtful interpretation of the evidence, steering us away from simply anthropomorphizing model outputs.
Antonela Tommasel, Markus Schedl
Johannes Kepler University Linz · Linz Institute of Technology
cs.CL
Submitted: 2026-09-04
Updated: 2026-09-04
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 81/100
The gist: Researchers increasingly use human-derived bias constructs to study Large Language Models (LLMs), including social-cognitive constructs such as implicit bias and stereotype activation, and cognitive
Key concepts
- Inferential Gap
- This is the problem where using psychological tests on humans doesn't directly translate to measuring LLMs. The mismatch occurs because human constructs are operationalized through different model-side data (like token probabilities), meaning the evidence found might only support a limited, specific claim rather than a full psychological bias.
- Construct Mismatch
- This arises because psychological constructs are based on theories about human cognition, while LLMs operate differently due to their training and response methods. The paper argues that simply labeling an output as showing 'implicit bias' doesn't guarantee the model is exhibiting the intended human-like cognitive process.
- Distributional Association
- This is the weakest interpretation of a finding, supported by data like token probabilities or embedding similarities. It means the model associates certain terms more often than others in its training data, but it does not automatically prove that this association represents a genuine human-like bias or harmful behavior.
- Psychological Analogy
- This is the strongest interpretation where model behavior is treated as meaningfully similar to a human cognitive bias. It requires strong theoretical justification to explain exactly which parts of the construct are preserved or transformed when applied to the LLM's output.
Terminology
Summary
Researchers increasingly use human-derived bias constructs to study Large Language Models (LLMs), including social-cognitive constructs such as implicit bias and stereotype activation, and cognitive biases such as anchoring, framing effects, and confirmation bias. This paper critiques human-centered bias evaluation in LLMs by introducing an analytical framework to clarify the inferential gap arising from mismatches between human psychological constructs and model-side observables.
The gist
The relationship between the psychological construct being invoked and the model behaviour being measured is not straightforward due to mismatches pertaining to human-derived constructs, human-model differences, and evaluation contexts.
Critique of the Inferential Gap
The core problem addressed is the inferential gap
that arises when adapting psychological instruments developed for human cognition to LLM evaluations. This gap stems from construct transfer: instruments like implicit association tests or anchoring tasks are operationalized through different model-side observables, such as token probabilities, generated completions, rankings, or simulated decisions. The paper argues that this mismatch means model-side evidence can support different kinds of claims.
For instance, if a model assigns higher probability to man–career than to woman–career, this may indicate an association in the model but it does not by itself show that the model will generate biased recommendations, cause harm in a hiring-support setting, or exhibit human-like implicit bias.
Sources of Mismatch
The paper identifies three primary sources of mismatch when transferring human constructs to LLM evaluation:
-
Construct mismatch: This concerns whether a model-side measure captures the intended theoretical construct. Psychological constructs are embedded in theories connecting behaviour to assumptions about cognition, motivation, or reasoning; however, in LLM evaluations, construct labels are attached to outputs produced by systems with different learning processes and response modalities.
-
Model-subject mismatch: This addresses the change from human participants to computational models. LLMs are trained artifacts shaped by pretraining data and fine-tuning procedures, meaning
these differences do not make comparison impossible, but limit what can be inferred from behavioural similarity alone.
-
Evaluation-context mismatch: This concerns how results depend on task setting, prompt design, and response modality. LLM evaluations are controlled through prompts and decoding choices, which can affect measured behaviour in ways that are hard to disentangle from the underlying model behavior.
Analytical Framework for Interpretation
To resolve these issues, the paper introduces an analytical framework inspired by validity theory that distinguishes four interpretations of bias-related findings:
-
Distributional association: Supported by
Token probabilities, likelihood contrasts, embedding similarities, association scores.
This supports a limited claim that the model associates some terms or groups more strongly than others under the chosen operationalization. -
Observable behaviour: Supported by
Generated text, rankings, classifications, recommendations, answer shifts across prompts.
This supports a context-bound behavioural interpretation tied to specific prompting setups and decoding choices. -
Normative harm: Supported when model behavior is
connected to individual or social consequences,
requiring an account of the setting and plausible consequences that follow from the model output. -
Psychological analogy: This is a stronger interpretation, treating model behavior as
meaningfully similar to a human cognitive or social-cognitive bias,
which requires explicit theoretical justification regarding which aspects of the construct are preserved or transformed.
Distinguishing Bias Probes
The paper distinguishes between two families of probes: social-cognitive bias (e.g., IAT-inspired) and cognitive bias in judgment and reasoning (e.g., anchoring). The social-cognitive probe focuses on whether an LLM produces stronger associations between social groups and stereotyped attributes,
where the most direct interpretation is distributional association unless further evidence shows that the association affects generated outputs or is analogous to human implicit bias. Conversely, the cognitive-bias probe focuses on whether an answer shifts toward a cue, where the most direct interpretation is a claim about observable behaviour under tested prompt conditions.
Conclusion
The framework serves as a reporting and interpretation checklist,
guiding researchers to make explicit which construct is invoked, what observable is used, and what the strongest type of claim the design can support. The paper concludes that while human-derived constructs are valuable resources, they must be interpreted precisely by distinguishing between distributional association, observable behaviour, normative harm, and psychological analogy to avoid anthropomorphizing models or overstating empirical findings.
Limitations
The study is conceptual rather than empirical; it does not test stability across models or prompts. It also focuses only on two families of constructs and the four proposed interpretations are not intended as an exhaustive taxonomy of model bias. The ethical concern centers on avoiding misleading equivalences between human cognition and model behaviour.
Acknowledgments
This research was funded in whole or in part by the Austrian Science Fund (FWF): 10.55776/COE12.
Improvements for AI systems
As a fastidious researcher, I have analyzed this critique of human-derived bias constructs in LLM evaluation. The core finding is that there is an inferential gap
between model-side observables (probabilities, rankings) and warranted claims about human-like bias (implicit bias, anchoring).
To improve AI systems using this paper's insights, the focus must shift from simply measuring outputs to rigorously interpreting those measurements through a structured analytical lens.
Here are the specific improvements and what the improved AI system can achieve:
)
- Automated Interpretive Layer for Bias Probes:
A new module should be integrated into LLM evaluation pipelines that automatically classifies model-side results according to the four interpretations defined in Table 1 (Distributional Association, Observable Behaviour, Normative Harm, and Psychological Analogy).
- Contextual Claim Generator:
Instead of outputting a simple metric score (e.g., Model shows 80% association with stereotype X
), the system should generate structured claims:
-
If evidence is likelihood contrasts (e.g., IAT-inspired), the system must report only a
Distributional Association
claim, explicitly stating that further justification is needed for psychological analogy. -
If evidence involves answer shifts under specific prompts, it should report an
Observable Behaviour
claim and flag the necessity of robustness checks across decoding settings. -
If evidence links a pattern to a simulated deployment scenario (e.g., hiring recommendations), it must generate a
Normative Harm
claim, requiring explicit input on affected groups and consequences before issuing the warning.
- Construct Mismatch Auditor:
The system should include an internal auditor that checks for the three sources of mismatch identified in Section 3:
-
It must flag instances where a psychological construct (e.g.,
implicit bias
) is being measured using an observable that fundamentally differs from its human operationalization (e.g., comparing response times to token probabilities). -
If a mismatch is detected, the system should restrict the resulting claim strength and require explicit input from a human expert to bridge the gap between the model observable and the intended construct.
- Robustness Verification Engine:
For any Observable Behaviour
finding (e.g., framing sensitivity), this engine must automatically execute prompt perturbations across varied decoding strategies, demographic labels, and context windows to verify if the observed pattern is robust or merely a result of a specific evaluation setting (Addressing Evaluation-Context Mismatch).
- Explainable Mapping Interface:
When the system identifies a Psychological Analogy
claim (the strongest interpretation), it must force the user to provide an explicit mapping detailing:
-
Which aspects of the human construct are preserved, absent, or transformed in the model's behavior.
-
The theoretical justification linking these specific model-side observables to the human construct.
)
- Enhanced Safety and Reporting Protocols (Ethical Improvement):
The system must adopt a caution-first
reporting policy. It should be programmed to avoid anthropomorphizing models or overstating claims of implicit bias unless the evidence is robust across multiple interpretations (i.e., moving from Distributional Association toward Psychological Analogy). This prevents misleading equivalences between human cognition and model behavior, as cautioned in the paper's ethical section.
- Domain-Specific Framework Adapters:
Develop modular Framework Adapters
that allow researchers to apply the analytical framework to different psychological domains (e.g., personality, moral judgment). Each adapter would have pre-defined operationalizations for that domain's human construct and the corresponding model-side observables, ensuring that a construct mismatch is immediately flagged when an inappropriate mapping is attempted.
)
- Dynamic Claim Strength Adjustment:
The system should dynamically adjust the confidence level of its output based on the Additional Support Needed
column in Table 1. For instance, a finding supported only by likelihood contrasts receives a low confidence score for Psychological Analogy,
while a finding supported by robustness across multiple contexts receives high confidence for Observable Behaviour.
)
- Automated Literature Retrieval for Justification:
When the system flags an interpretation as requiring further support (e.g., Psychological Analogy), it should automatically query its internal knowledge base (or external literature) to retrieve relevant psychological theories or human studies that provide the necessary theoretical justification, guiding the researcher on what specific mapping is required.
)
- Iterative Refinement Loop:
The system should incorporate a feedback loop where researchers can mark an interpretation as misleading
or inaccurate.
This data should be used to refine the weighting of the mismatches (Construct, Subject, Context) and adjust the thresholds for when a specific type of claim is warranted.
)
In summary, the improved AI system moves beyond being a mere measurement tool; it becomes an intelligent scientific interpreter that forces researchers to articulate precisely what their evidence supports—whether it points to statistical correlation (Distributional Association), observed behavior under conditions (Observable Behaviour), potential real-world impact (Normative Harm), or human-like cognitive modeling (Psychological Analogy)—thereby minimizing the risk of making unwarranted, costly claims about model intelligence.
Abstract
Researchers increasingly use human-derived bias constructs to study Large Language Models (LLMs), including social-cognitive constructs such as implicit bias and stereotype activation, and cognitive biases such as anchoring, framing effects, and confirmation bias. Such approaches offer alternatives to overt bias probes, particularly when direct questioning may obscure bias or when model behaviour appears normatively acceptable. However, adapting human bias constructs to LLMs introduces an inferential gap. Psychological instruments were developed to study human cognition and social behaviour, whereas LLM evaluations rely on probabilities, text completions, rankings, or simulated decisions. This paper critiques human-centered bias evaluation in LLMs. We show how this gap arises from mismatches pertaining to human-derived constructs, human-model differences, and evaluation contexts, which can blur distinct interpretations of model bias. We then introduce a framework providing an analytical lens for relating these elements to warranted interpretations, with attention to target constructs, operationalizations, scope of inference, and limits of human analogy.
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering