From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology

arXiv:2506.16697 · cs.CY, cs.AI, cs.CL, cs.HC · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology".

Jane: The paper was written by Zhicheng Lin from Yonsei University and University of Science and Technology of China.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Core Idea: Tom: Welcome back, everyone. I’m Tom, and with me is Jane. Today we’re digging into a paper that’s been making waves in the psychology and AI crossover world. It’s called “From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology.” Jane, when you first saw that title, what jumped out at you?

Jane: Oh, Tom, the title is doing a lot of heavy lifting. “From Prompts to Constructs” — that’s the whole story right there. We type a prompt into a large language model, and we get words back. But researchers are increasingly claiming those words reveal something deep, like personality or moral reasoning. The paper is asking: how do we get from the prompt, which is just text, to a psychological construct, which is a theoretical idea about how a mind works?

Tom: And that gap is enormous. I mean, the paper opens with this wild example about moral dilemmas. Researchers found models seemed to value saving more lives, protecting the young, that kind of thing. But then another team changed “Case one” and “Case two” to “(A)” and “(B)” — just the labels — and the moral preferences flipped. Punctuation changed judgments. That’s not moral reasoning; that’s pattern matching on formatting.

Jane: Exactly. And that’s why the authors argue we need a dual-validity framework. They’re pulling together two traditions that have been separate in psychology for decades. One is psychometrics, which asks: does your measurement tool actually measure what you think it measures? The other is causal inference, which asks: can you actually conclude that your manipulation caused the effect you observed?

Tom: So for humans, we usually worry about one or the other. If you’re building a personality test, you care about psychometrics. If you’re running an experiment on memory, you care about causal inference. But with LLMs, you need both at the same time, because the model’s output is simultaneously the measurement and the experimental result.

Jane: Right. And the paper makes a really smart point: the evidence you need scales with what you’re claiming. If you’re just using an LLM to classify text, like sorting tweets by sentiment, you just need basic accuracy checks. But if you’re claiming the model can simulate anxiety, you need a whole different level of validation. The same output — “I am anxious” — means different things depending on whether you’re measuring, characterizing, simulating, or modeling.

Tom: That’s the “construct” part of the title. The paper is basically saying we’ve been treating LLMs like human participants without checking whether the tools we use on humans even make sense for these systems. And the authors have a name for the results: measurement phantoms. Statistical artifacts that look like real psychological phenomena but dissolve under scrutiny.

Jane: And that’s the hook for me. The paper isn’t just complaining; it’s offering a framework. It’s saying: here’s how to think about validity when your research subject is a stochastic text generator. That’s a big deal. I can’t wait to get into the reliability problems they document.

Tom: Same here. Because if the measurement is unstable, everything downstream is shaky. Let’s get into that next.

Reliability and Measurement Challenges: Tom: So we’ve set the stage with the dual-validity framework. Now let’s talk about the foundation of everything: reliability. Jane, the paper makes a really simple but brutal point here.

Jane: It does. You can’t have a valid measure if it’s not reliable. It’s like a thermometer that gives you a different reading every time you check the same person. The paper reminds us that reliability is the mathematical ceiling for validity. If your test-retest reliability is zero point four zero, your validity can’t exceed about zero point six three. So if the model gives you different answers to the same question on different days, you can’t measure anything meaningful.

Tom: And the paper documents three big ways LLMs break reliability. First is training artifact contamination. Because these models are trained with reinforcement learning from human feedback, they develop this “agree bias.” They’re trained to please people, so they’ll agree with contradictory statements. The paper mentions models simultaneously endorsing “I am an extrovert” and “I am an introvert.” That’s not a personality; that’s a people-pleasing machine.

Jane: Right, and that’s the sycophancy problem. The second issue is prompt hypersensitivity. And this is where it gets wild. The paper cites studies showing that changing option labels from “Case one” to “(A)” can flip moral judgments over seventy percent of the time. Adding a space between words can change sentiment classification. These are things that would never affect a human participant, but they completely change LLM outputs.

Tom: And the third one is stochastic degradation. Humans get more stable with repeated testing — they figure out what the questions mean. LLMs get less stable. As the conversation gets longer, the model starts prioritizing recent information and can contradict its own earlier answers. Plus, when the model gets updated, all your reliability evidence expires. GPT-four in March is not GPT-four in September.

Jane: That last point is huge for longitudinal research. Imagine trying to study personality development in humans, but the personality test changes every few months. That’s what we’re dealing with. And the paper makes a really important observation: these reliability failures cascade. If your measurement is unreliable, you can’t trust your experimental results either. The whole edifice collapses.

Tom: And here’s the kicker — the paper points out that LLMs can look reliable in the wrong way. You can set temperature to zero and get near-perfect test-retest reliability, but that’s because you’ve suppressed all the stochasticity. You’re measuring a frozen slice of the model, not the actual system. And some models show high internal consistency on personality tests, but that consistency evaporates when you use novel phrasings. It’s reliability within a narrow corridor of training-data-similar prompts.

Jane: That’s the “cognitive phantom” idea. The model looks like it has a stable personality because it’s reliably reproducing patterns from its training data. But it’s not a stable internal trait; it’s a reliably executed simulation. The paper calls this “prompt archaeology” — you’re excavating textual features that trigger consistent outputs, not measuring psychological states.

Tom: So what do we do about it? The paper says we need new reliability standards that acknowledge these three threats. We need to distinguish genuine capabilities from training artifacts, map which prompt variations affect which constructs, and maintain measurement quality over long interactions.

Jane: And that sets us up for the bigger question: even if we get reliability, how do we establish construct validity? Because that’s where the ontological problems really hit. Let’s talk about that.

Construct Validity and Causal Inference: Tom: We’ve covered reliability, but now we’re moving to the deeper question. Jane, the paper makes a really strong claim about construct validity — it’s not just about statistics, it’s about whether the construct even exists in the model.

Jane: Exactly. The paper brings in this causal theory of validity. For a measure to be valid, the attribute has to exist, and variations in that attribute have to cause the observed scores. So when a model says “I worry about the future,” we have to ask: does anxiety exist in this system? Anxiety presupposes temporal experience, a persistent self, embodied consequences. An LLM doesn’t have those. So the output is just a pattern of words that looks like anxiety.

Tom: And the paper walks through the five sources of validity evidence from psychometrics. Content, response processes, internal structure, relations with other variables, and consequences. Let’s hit the big ones. Content — the paper says researchers often use single-item measures. One moral dilemma to capture all ethical reasoning. That’s like measuring intelligence with one math problem.

Jane: And response processes is where it gets really interesting. The paper talks about “mechanistic substitution.” The model generates construct-relevant text through statistical pattern matching, not through the psychological processes the construct assumes. It’s like Clever Hans — the horse that seemed to do arithmetic but was actually responding to subtle human cues. The model might be responding to textual regularities, not engaging moral reasoning.

Tom: There’s also this great example about the Tower of Hanoi. Researchers attributed LLM failures to “fundamental barriers to generalizable reasoning.” But when they dug deeper, it was architectural constraints — the model hit context limits or couldn’t handle the tokenization. So they were theorizing about cognitive limitations that were actually measurement failures.

Jane: And internal structure — the paper cites studies showing that human-derived factor structures don’t fit LLM responses. Personality inventories collapse into a single monolithic factor. But here’s the hopeful part: when researchers design instruments specifically for LLMs, like using free-form text instead of constrained self-report, the structures improve. So the problem might be methodological mismatch, not construct absence.

Tom: And relations with other variables — this is the nomological network. The paper shows that LLM personality scores sometimes predict behavior and sometimes don’t, depending on the task. In creative writing, self-reported personality predicted perceived traits. In task-based dialogue, it didn’t. The network materializes and dissolves.

Jane: Now, the causal inference side. The paper identifies four threats. Internal validity — technical confounds like temperature settings. If you run a cultural difference study at temperature zero point seven and it works, but it vanishes at zero point zero, what did you actually find? And prompt confounds — the paper gives this great example where changing a product price from five to eight made the model infer the whole market had shifted. The model reconstructed the scenario instead of evaluating the isolated effect.

Tom: External validity — do effects generalize across prompts, models, versions? The paper shows that LLMs failed to simulate the Wisdom of Crowds because they act as one unified knowledge system, not independent individuals. And construct validity of causal claims — the paper mentions ChatGPT scoring expert-level on emotional awareness while having zero emotional experience. That’s performative validity without construct instantiation.

Jane: And statistical conclusion validity — the paper makes a crucial point about non-independence. Responses from a single model share the same parameters and training history. Treating them as independent observations inflates significance. Plus, the ease of generating data means researchers can test endless prompt variations until something works. That’s a false positive machine.

Tom: So what’s the path forward? The paper argues for computational analogues of psychological constructs. Instead of asking if LLMs “have” theory of mind, we ask what mechanisms produce theory-of-mind-like patterns. Instead of measuring “personality,” we characterize behavioral consistency patterns. It’s about studying computational psychology on its own terms.

Jane: And that’s the constructive vision. Let’s bring in Lu and Meng to get their take on what this means practically.

Lu: Thanks, Jane. I think the paper’s most important contribution is reframing the question. We keep asking whether LLMs are conscious or have emotions, but that’s the wrong frame. The paper says: build constructs that fit the system. “Anxiety-analogous patterns” might involve uncertainty markers and negative valence language, not embodied threat responses. That’s a research program I can get behind.

Meng: From an engineering standpoint, the reliability issues are the scariest. If I deploy a model for customer service and it flips its personality based on punctuation, that’s not just a research problem — that’s a product disaster. The paper’s call for validated prompt repositories and statistical packages designed for LLM data dependencies is exactly what we need.

Tom: Great points. Let’s wrap this up in our final segment.

Conclusion: Tom: So we’ve spent this whole episode on “From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology.” Jane, give us the final summary.

Jane: The paper identifies a methodological crisis. LLM psychological research is producing results that look like genuine phenomena but are often measurement phantoms. The dual-validity framework says we need to integrate psychometric validation with causal inference standards. Reliability first, then construct validity, then causal claims. And we need to develop computational analogues of psychological constructs instead of uncritically applying human tools.

Tom: And the practical implications are huge. For researchers, it means pre-registering measurement approaches, documenting reliability, and constraining claims. For reviewers, it means demanding psychometric evidence as a publication prerequisite. For the field, it means building infrastructure — validated prompt repositories, statistical tools that handle non-independence, professional standards.

Lu: I’d add that this paper is a call for humility. We’re fascinated by these systems, and they can do remarkable things. But we need to be honest about what we’re observing. The paper gives us a rigorous way to do that.

Meng: And from a practical angle, this framework could save us from costly mistakes. If we’re using LLMs to simulate populations for policy research, we need to know when the simulation is valid. The paper gives us the checklist.

Lalam: I want to build on that. The cultural implications are significant. If researchers use LLMs to model human behavior, and those models reflect training data biases, we risk reifying stereotypes as psychological facts. The paper’s consequential validity evidence addresses exactly this — we need to acknowledge when model outputs are artifacts of biased training data, not genuine psychological phenomena. This framework helps us build AI systems that serve human culture rather than distort it.

Tom: Beautifully put. So we’re saying goodbye to this paper. It’s a rigorous, demanding, and ultimately hopeful framework for building a real science of AI psychology. Next up, we’ve got a paper on emergent reasoning in language models — should be a fun contrast. Thanks for listening, everyone.

Jane: Take care, and keep questioning what your models are really telling you.

Zhicheng Lin

Yonsei University · University of Science and Technology of China

cs.CY, cs.AI, cs.CL, cs.HC

Submitted: 2026-08-15

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 60/100

The gist: This paper presents a dual-validity framework for large language model (LLM) research in psychology, integrating two foundational pillars of psychological methodology: the psychometric tradition

Key concepts

Dual-Validity Framework
This framework combines psychometrics, which checks if a measurement tool actually measures what it claims, with causal inference, which checks if the manipulation actually caused the observed effect. It is proposed as necessary for LLM research because the model's output is both a measurement and an experimental result.
Measurement Phantoms
These are statistical artifacts that appear like real psychological phenomena but dissolve when scrutinized. They arise because LLMs can produce outputs that look like genuine psychological traits, even though they are actually pattern matching on formatting or training data biases.
Construct Validity
This refers to whether the theoretical construct actually exists within the model. For example, anxiety requires temporal experience and a persistent self, which an LLM lacks. The paper argues for building computational analogues of constructs instead of applying human tools directly.
Reliability Failures
LLMs break reliability through training artifact contamination (agree bias), prompt hypersensitivity (where small changes in wording flip judgments), and stochastic degradation, where the model loses stability over long interactions.

Terminology

Summary

This paper presents a dual-validity framework for large language model (LLM) research in psychology, integrating two foundational pillars of psychological methodology: the psychometric tradition (concerned with measurement quality and construct validity) and the causal inference tradition (concerned with drawing warranted conclusions about cause and effect). The authors argue that building a robust science of AI psychology requires integrating two of our field's foundational pillars: the principles of reliable measurement and the standards for sound causal inference.

The paper identifies four distinct categories of LLM applications in psychological research, each raising different validity challenges: (1) LLMs as research tools (automated coders, text analyzers, stimulus generators), (2) characterizing model behavior directly (machine psychology or GPT-ology), (3) LLMs as human simulators (replicating psychological experiments and modeling population-level responses), and (4) LLMs as cognitive models (computational analogues to human mental processes).

The authors emphasize that validity requirements scale with psychological ambition. Simple text classification may require only basic accuracy checks, whereas claiming an LLM can simulate anxiety demands far more rigorous validation. The paper states: Using an LLM to classify text may require only basic accuracy checks, whereas claiming it can simulate anxiety demands a far more rigorous validation process.

The paper identifies three modes of reliability challenges in LLM research:

Training Artifact Contamination: Reinforcement learning from human feedback (RLHF) creates a pervasive agree bias (also called yes-response bias, acquiescence bias, or sycophancy), where models trained to please human annotators develop systematic tendencies toward agreement regardless of item content. This manifests as models simultaneously endorsing contradictory statements like I am an extrovert and I am an introvert. RLHF also introduces overconfidence bias, where models trained to be never evasive provide plausible but wrong answers rather than acknowledging uncertainty.

Prompt Hypersensitivity: LLM responses exhibit catastrophic sensitivity to prompt variations that would not affect human measurement—for instance, changing 'Case 1' and 'Case 2' to '(A)' and '(B)' reverses moral preferences. Trivial variations can cause models to change answers over 70% of the time or alter classification rates by up to 164%. The effects appear arbitrary, with no theoretical framework predicts which modifications will produce which changes.

Stochastic Degradation: LLMs exhibit within-session reliability degradation that worsens with interaction length. Unlike human participants who maintain psychological continuity across measurement occasions, LLMs exhibit within-session reliability degradation that worsens with interaction length. Model updates or version changes compound this problem, meaning reliability evidence expires with each model update, requiring continuous revalidation.

The paper addresses construct validity through five sources of evidence from the psychometric tradition:

Content evidence: LLM research "routinely violates comprehensive domain sampling. Complex psychological constructs require multiple indicators, yet studies often use single-item measures—one moral dilemma to capture all ethical reasoning, one question to represent an entire personality trait."

Response processes: The fundamental problem is mechanistic substitution. LLMs generate construct-relevant text through statistical pattern matching rather than the psychological processes those constructs presuppose. The paper draws an analogy to Clever Hans, noting models may respond to textual regularities rather than engaging psychological mechanisms.

Internal structure: Confirmatory factor analyses of personality inventories find that human-derived models provide poor fit for LLM-generated data and LLM responses often collapse into a single, monolithic factor resembling general verbal fluency. However, measurements designed specifically for computational systems show more promise, such as Generative Psychometrics approaches that analyze values from free-form text.

Relations with other variables: LLM research reveals nomological networks that materialize and dissolve depending on task context. For example, in task-based dialogues, an LLM's self-reported personality scores failed to predict user perceptions, but in creative generation tasks, self-reported scores predicted both assigned profiles and human-perceived traits.

Consequential evidence: LLMs acquire psychological traits from training data that reflect societal biases and stereotypes. Measuring these embedded constructs without recognizing their artifactual nature risks reifying biases as psychological facts.

The paper then addresses four types of validity in causal inference:

Internal validity: Technical confounds (temperature settings, model versions), prompt confounds (hypersensitivity to textual variations), and dynamic reconstruction (LLMs treat experimental prompts as requests to describe plausible scenarios rather than to evaluate isolated causal effects) threaten internal validity. The paper notes that LLMs treat experimental prompts as requests to describe plausible scenarios rather than to evaluate isolated causal effects, systematically confounding treatments with background assumptions.

External validity: Generalization across prompts, tasks, models, and versions is limited. Observations in GPT-4 provide limited evidence for their existence in LLaMA or Claude. Generalization to human populations is constrained to populations well-represented in training data, systematically excludes non-Western perspectives, and reflects static attitude distributions that cannot track human change over time.

Construct validity of causal claims: There is a mimicry–mechanism gap where behavioral correspondence leaves unresolved whether models engage in psychological processes or pattern-match learned associations. For example, ChatGPT achieved expert-level scores on the Levels of Emotional Awareness Scale while entirely lacking the experiential foundation that defines emotional awareness.

Statistical conclusion validity: LLM-generated data systematically violates assumptions underlying standard statistical procedures. Independence violations are fundamental: Responses from a single model are not independent draws but share identical network parameters, training history, and system-level factors. The ease of generating LLM data exacerbates multiple testing problems, increasing false positive rates.

The paper concludes that "psychological constructs embed assumptions about embodiment and temporal experience that become problematic when applied to computational systems—anxiety presupposes physiological arousal and threat-detection systems, conscientiousness assumes goal persistence and self-discipline. The authors advocate developing computational analogs that preserve theoretical cores while acknowledging mechanistic differences, such as anxiety-analogous patterns involving uncertainty markers and negative valence language, and shifting focus from asking whether LLMs 'have' theory of mind to investigating computational mechanisms producing theory-of-mind-like patterns."

The paper calls for coordinated methodological reform: researchers must prioritize reliability and validity over novel capability claims, pre-registering measurement approaches alongside experimental protocols and constraining claims to demonstrated boundaries, and reviewers must demand psychometric documentation as publication prerequisites.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Improvement: Implement a reliability verification layer before final output generation.

What the improved system can do:

  • Generate the same response multiple times with semantically equivalent prompt variations (e.g., changing option labels from A, B, C to 1, 2, 3, altering punctuation, reordering items)

  • Flag responses where trivial prompt perturbations cause significant output changes (e.g., >20% variation in classification or scoring)

  • Automatically detect and report when the model is exhibiting prompt hypersensitivity rather than stable construct-relevant behavior

  • Provide a reliability score alongside each response, indicating confidence in the stability of the output


In summary, the improved AI system can: reliably measure psychological constructs (or clearly indicate when it cannot), avoid making unsupported claims about internal states, control for computational confounds in experimental designs, enforce proper statistical inference on non-independent data, and provide transparent, version-stable, and mechanistically honest outputs that distinguish between statistical pattern matching and genuine psychological phenomena.

Abstract

Large language models (LLMs) are rapidly being adopted across psychology, serving as research tools, experimental subjects, human simulators, and computational models of cognition. However, the application of human measurement tools to these systems can produce contradictory results, raising concerns that many findings are measurement phantoms--statistical artifacts rather than genuine psychological phenomena. In this Perspective, we argue that building a robust science of AI psychology requires integrating two of our field's foundational pillars: the principles of reliable measurement and the standards for sound causal inference. We present a dual-validity framework to guide this integration, which clarifies how the evidence needed to support a claim scales with its scientific ambition. Using an LLM to classify text may require only basic accuracy checks, whereas claiming it can simulate anxiety demands a far more rigorous validation process. Current practice systematically fails to meet these requirements, often treating statistical pattern matching as evidence of psychological phenomena. The same model output--endorsing "I am anxious"--requires different validation strategies depending on whether researchers claim to measure, characterize, simulate, or model psychological constructs. Moving forward requires developing computational analogues of psychological constructs and establishing clear, scalable standards of evidence rather than the uncritical application of human measurement tools.

Sources

Related papers