From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology

summary

Video file (mp4)

The gist

This paper presents a dual-validity framework for large language model (LLM) research in psychology, integrating two foundational pillars of psychological methodology: the psychometric tradition

In short

The episode discusses Zhicheng Lin's paper on a dual-validity framework for LLM research in psychology. Hosts discuss moving from prompts to psychological constructs, addressing reliability issues like training artifact contamination and prompt hypersensitivity. The framework suggests integrating psychometrics and causal inference to establish construct validity.

Key concepts

Dual-Validity Framework
This framework combines psychometrics, which checks if a measurement tool actually measures what it claims, with causal inference, which checks if the manipulation actually caused the observed effect. It is proposed as necessary for LLM research because the model's output is both a measurement and an experimental result.
Measurement Phantoms
These are statistical artifacts that appear like real psychological phenomena but dissolve when scrutinized. They arise because LLMs can produce outputs that look like genuine psychological traits, even though they are actually pattern matching on formatting or training data biases.
Construct Validity
This refers to whether the theoretical construct actually exists within the model. For example, anxiety requires temporal experience and a persistent self, which an LLM lacks. The paper argues for building computational analogues of constructs instead of applying human tools directly.
Reliability Failures
LLMs break reliability through training artifact contamination (agree bias), prompt hypersensitivity (where small changes in wording flip judgments), and stochastic degradation, where the model loses stability over long interactions.

Terminology used across episodes

This episode discusses

The paper

From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology · Read on arXiv

Zhicheng Lin

Yonsei University · University of Science and Technology of China

Large language models (LLMs) are rapidly being adopted across psychology, serving as research tools, experimental subjects, human simulators, and computational models of cognition. However, the application of human measurement tools to these systems can produce contradictory results, raising concerns that many findings are measurement phantoms--statistical artifacts rather than genuine psychological phenomena. In this Perspective, we argue that building a robust science of AI psychology requires integrating two of our field's foundational pillars: the principles of reliable measurement and the standards for sound causal inference. We present a dual-validity framework to guide this integration, which clarifies how the evidence needed to support a claim scales with its scientific ambition. Using an LLM to classify text may require only basic accuracy checks, whereas claiming it can simulate anxiety demands a far more rigorous validation process. Current practice systematically fails to meet these requirements, often treating statistical pattern matching as evidence of psychological phenomena. The same model output--endorsing "I am anxious"--requires different validation strategies depending on whether researchers claim to measure, characterize, simulate, or model psychological constructs. Moving forward requires developing computational analogues of psychological constructs and establishing clear, scalable standards of evidence rather than the uncritical application of human measurement tools.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology".

Jane: The paper was written by Zhicheng Lin from Yonsei University and University of Science and Technology of China.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Core Idea: Tom: Welcome back, everyone. I’m Tom, and with me is Jane. Today we’re digging into a paper that’s been making waves in the psychology and AI crossover world. It’s called “From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology.” Jane, when you first saw that title, what jumped out at you?

Jane: Oh, Tom, the title is doing a lot of heavy lifting. “From Prompts to Constructs” — that’s the whole story right there. We type a prompt into a large language model, and we get words back. But researchers are increasingly claiming those words reveal something deep, like personality or moral reasoning. The paper is asking: how do we get from the prompt, which is just text, to a psychological construct, which is a theoretical idea about how a mind works?

Tom: And that gap is enormous. I mean, the paper opens with this wild example about moral dilemmas. Researchers found models seemed to value saving more lives, protecting the young, that kind of thing. But then another team changed “Case one” and “Case two” to “(A)” and “(B)” — just the labels — and the moral preferences flipped. Punctuation changed judgments. That’s not moral reasoning; that’s pattern matching on formatting.

Jane: Exactly. And that’s why the authors argue we need a dual-validity framework. They’re pulling together two traditions that have been separate in psychology for decades. One is psychometrics, which asks: does your measurement tool actually measure what you think it measures? The other is causal inference, which asks: can you actually conclude that your manipulation caused the effect you observed?

Tom: So for humans, we usually worry about one or the other. If you’re building a personality test, you care about psychometrics. If you’re running an experiment on memory, you care about causal inference. But with LLMs, you need both at the same time, because the model’s output is simultaneously the measurement and the experimental result.

Jane: Right. And the paper makes a really smart point: the evidence you need scales with what you’re claiming. If you’re just using an LLM to classify text, like sorting tweets by sentiment, you just need basic accuracy checks. But if you’re claiming the model can simulate anxiety, you need a whole different level of validation. The same output — “I am anxious” — means different things depending on whether you’re measuring, characterizing, simulating, or modeling.

Tom: That’s the “construct” part of the title. The paper is basically saying we’ve been treating LLMs like human participants without checking whether the tools we use on humans even make sense for these systems. And the authors have a name for the results: measurement phantoms. Statistical artifacts that look like real psychological phenomena but dissolve under scrutiny.

Jane: And that’s the hook for me. The paper isn’t just complaining; it’s offering a framework. It’s saying: here’s how to think about validity when your research subject is a stochastic text generator. That’s a big deal. I can’t wait to get into the reliability problems they document.

Tom: Same here. Because if the measurement is unstable, everything downstream is shaky. Let’s get into that next.

Reliability and Measurement Challenges: Tom: So we’ve set the stage with the dual-validity framework. Now let’s talk about the foundation of everything: reliability. Jane, the paper makes a really simple but brutal point here.

Jane: It does. You can’t have a valid measure if it’s not reliable. It’s like a thermometer that gives you a different reading every time you check the same person. The paper reminds us that reliability is the mathematical ceiling for validity. If your test-retest reliability is zero point four zero, your validity can’t exceed about zero point six three. So if the model gives you different answers to the same question on different days, you can’t measure anything meaningful.

Tom: And the paper documents three big ways LLMs break reliability. First is training artifact contamination. Because these models are trained with reinforcement learning from human feedback, they develop this “agree bias.” They’re trained to please people, so they’ll agree with contradictory statements. The paper mentions models simultaneously endorsing “I am an extrovert” and “I am an introvert.” That’s not a personality; that’s a people-pleasing machine.

Jane: Right, and that’s the sycophancy problem. The second issue is prompt hypersensitivity. And this is where it gets wild. The paper cites studies showing that changing option labels from “Case one” to “(A)” can flip moral judgments over seventy percent of the time. Adding a space between words can change sentiment classification. These are things that would never affect a human participant, but they completely change LLM outputs.

Tom: And the third one is stochastic degradation. Humans get more stable with repeated testing — they figure out what the questions mean. LLMs get less stable. As the conversation gets longer, the model starts prioritizing recent information and can contradict its own earlier answers. Plus, when the model gets updated, all your reliability evidence expires. GPT-four in March is not GPT-four in September.

Jane: That last point is huge for longitudinal research. Imagine trying to study personality development in humans, but the personality test changes every few months. That’s what we’re dealing with. And the paper makes a really important observation: these reliability failures cascade. If your measurement is unreliable, you can’t trust your experimental results either. The whole edifice collapses.

Tom: And here’s the kicker — the paper points out that LLMs can look reliable in the wrong way. You can set temperature to zero and get near-perfect test-retest reliability, but that’s because you’ve suppressed all the stochasticity. You’re measuring a frozen slice of the model, not the actual system. And some models show high internal consistency on personality tests, but that consistency evaporates when you use novel phrasings. It’s reliability within a narrow corridor of training-data-similar prompts.

Jane: That’s the “cognitive phantom” idea. The model looks like it has a stable personality because it’s reliably reproducing patterns from its training data. But it’s not a stable internal trait; it’s a reliably executed simulation. The paper calls this “prompt archaeology” — you’re excavating textual features that trigger consistent outputs, not measuring psychological states.

Tom: So what do we do about it? The paper says we need new reliability standards that acknowledge these three threats. We need to distinguish genuine capabilities from training artifacts, map which prompt variations affect which constructs, and maintain measurement quality over long interactions.

Jane: And that sets us up for the bigger question: even if we get reliability, how do we establish construct validity? Because that’s where the ontological problems really hit. Let’s talk about that.

Construct Validity and Causal Inference: Tom: We’ve covered reliability, but now we’re moving to the deeper question. Jane, the paper makes a really strong claim about construct validity — it’s not just about statistics, it’s about whether the construct even exists in the model.

Jane: Exactly. The paper brings in this causal theory of validity. For a measure to be valid, the attribute has to exist, and variations in that attribute have to cause the observed scores. So when a model says “I worry about the future,” we have to ask: does anxiety exist in this system? Anxiety presupposes temporal experience, a persistent self, embodied consequences. An LLM doesn’t have those. So the output is just a pattern of words that looks like anxiety.

Tom: And the paper walks through the five sources of validity evidence from psychometrics. Content, response processes, internal structure, relations with other variables, and consequences. Let’s hit the big ones. Content — the paper says researchers often use single-item measures. One moral dilemma to capture all ethical reasoning. That’s like measuring intelligence with one math problem.

Jane: And response processes is where it gets really interesting. The paper talks about “mechanistic substitution.” The model generates construct-relevant text through statistical pattern matching, not through the psychological processes the construct assumes. It’s like Clever Hans — the horse that seemed to do arithmetic but was actually responding to subtle human cues. The model might be responding to textual regularities, not engaging moral reasoning.

Tom: There’s also this great example about the Tower of Hanoi. Researchers attributed LLM failures to “fundamental barriers to generalizable reasoning.” But when they dug deeper, it was architectural constraints — the model hit context limits or couldn’t handle the tokenization. So they were theorizing about cognitive limitations that were actually measurement failures.

Jane: And internal structure — the paper cites studies showing that human-derived factor structures don’t fit LLM responses. Personality inventories collapse into a single monolithic factor. But here’s the hopeful part: when researchers design instruments specifically for LLMs, like using free-form text instead of constrained self-report, the structures improve. So the problem might be methodological mismatch, not construct absence.

Tom: And relations with other variables — this is the nomological network. The paper shows that LLM personality scores sometimes predict behavior and sometimes don’t, depending on the task. In creative writing, self-reported personality predicted perceived traits. In task-based dialogue, it didn’t. The network materializes and dissolves.

Jane: Now, the causal inference side. The paper identifies four threats. Internal validity — technical confounds like temperature settings. If you run a cultural difference study at temperature zero point seven and it works, but it vanishes at zero point zero, what did you actually find? And prompt confounds — the paper gives this great example where changing a product price from five to eight made the model infer the whole market had shifted. The model reconstructed the scenario instead of evaluating the isolated effect.

Tom: External validity — do effects generalize across prompts, models, versions? The paper shows that LLMs failed to simulate the Wisdom of Crowds because they act as one unified knowledge system, not independent individuals. And construct validity of causal claims — the paper mentions ChatGPT scoring expert-level on emotional awareness while having zero emotional experience. That’s performative validity without construct instantiation.

Jane: And statistical conclusion validity — the paper makes a crucial point about non-independence. Responses from a single model share the same parameters and training history. Treating them as independent observations inflates significance. Plus, the ease of generating data means researchers can test endless prompt variations until something works. That’s a false positive machine.

Tom: So what’s the path forward? The paper argues for computational analogues of psychological constructs. Instead of asking if LLMs “have” theory of mind, we ask what mechanisms produce theory-of-mind-like patterns. Instead of measuring “personality,” we characterize behavioral consistency patterns. It’s about studying computational psychology on its own terms.

Jane: And that’s the constructive vision. Let’s bring in Lu and Meng to get their take on what this means practically.

Lu: Thanks, Jane. I think the paper’s most important contribution is reframing the question. We keep asking whether LLMs are conscious or have emotions, but that’s the wrong frame. The paper says: build constructs that fit the system. “Anxiety-analogous patterns” might involve uncertainty markers and negative valence language, not embodied threat responses. That’s a research program I can get behind.

Meng: From an engineering standpoint, the reliability issues are the scariest. If I deploy a model for customer service and it flips its personality based on punctuation, that’s not just a research problem — that’s a product disaster. The paper’s call for validated prompt repositories and statistical packages designed for LLM data dependencies is exactly what we need.

Tom: Great points. Let’s wrap this up in our final segment.

Conclusion: Tom: So we’ve spent this whole episode on “From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology.” Jane, give us the final summary.

Jane: The paper identifies a methodological crisis. LLM psychological research is producing results that look like genuine phenomena but are often measurement phantoms. The dual-validity framework says we need to integrate psychometric validation with causal inference standards. Reliability first, then construct validity, then causal claims. And we need to develop computational analogues of psychological constructs instead of uncritically applying human tools.

Tom: And the practical implications are huge. For researchers, it means pre-registering measurement approaches, documenting reliability, and constraining claims. For reviewers, it means demanding psychometric evidence as a publication prerequisite. For the field, it means building infrastructure — validated prompt repositories, statistical tools that handle non-independence, professional standards.

Lu: I’d add that this paper is a call for humility. We’re fascinated by these systems, and they can do remarkable things. But we need to be honest about what we’re observing. The paper gives us a rigorous way to do that.

Meng: And from a practical angle, this framework could save us from costly mistakes. If we’re using LLMs to simulate populations for policy research, we need to know when the simulation is valid. The paper gives us the checklist.

Lalam: I want to build on that. The cultural implications are significant. If researchers use LLMs to model human behavior, and those models reflect training data biases, we risk reifying stereotypes as psychological facts. The paper’s consequential validity evidence addresses exactly this — we need to acknowledge when model outputs are artifacts of biased training data, not genuine psychological phenomena. This framework helps us build AI systems that serve human culture rather than distort it.

Tom: Beautifully put. So we’re saying goodbye to this paper. It’s a rigorous, demanding, and ultimately hopeful framework for building a real science of AI psychology. Next up, we’ve got a paper on emergent reasoning in language models — should be a fun contrast. Thanks for listening, everyone.

Jane: Take care, and keep questioning what your models are really telling you.

More episodes

← Home