A validity-guided workflow for robust large language model research in psychology
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "A validity-guided workflow for robust large language model research in psychology".
Jane: The paper was written by Zhicheng Lin from Yonsei University and University of Science and Technology of China.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the channel, everyone. I'm Tom, and with me as always is Jane. Today we're digging into a paper that's been making waves in the methods community, and it's called "A validity-guided workflow for robust large language model research in psychology."
Jane: Tom, I have to say, when I first saw this title, I thought, okay, another paper telling us to be careful with AI. But this one is actually different. It's by Zhicheng Lin, and it's basically a field guide for anyone who wants to use large language models in psychological research without tripping over their own feet.
Tom: And that's a real problem right now, right? Because there's this flood of papers claiming that ChatGPT has a personality, or that it has theory of mind, or that it's anxious. And then the next paper comes along and shows those results vanish if you change a single word in the prompt.
Jane: Exactly. And the paper has this great term for that. It calls those results "measurement phantoms." Statistical artifacts that look like real psychological phenomena but are really just the model reacting to formatting quirks or training data patterns.
Tom: Measurement phantoms. That's such a vivid way to put it. So the author's argument is that psychology already has a rigorous toolkit for making sure your measurements actually measure what you think they measure. It's called psychometrics. And we should be applying that toolkit to AI models instead of just assuming they work like human participants.
Jane: Right, and the paper's big move is to take that psychometric tradition and combine it with causal inference standards. So you're not just asking, does this scale hang together statistically? You're also asking, when I manipulate something, can I actually attribute the change to my manipulation and not to some computational confound?
Tom: And that's the dual-validity framework. I love that the paper doesn't just say, here's a problem. It actually gives you a six-stage workflow to solve it. We're going to walk through those stages in a bit, but first, Jane, what struck you most about the framing?
Jane: I think the most striking part is the opening examples. The paper cites a study where GPT-four's accuracy on a false-belief task went from ninety-seven point five percent to zero percent just by changing "in a box" to "on a box." That's not a subtle difference. That's the model completely flipping its answer based on a preposition.
Tom: Zero percent. That's wild. And there's another example where models endorse both "I am an introvert" and "I am an extrovert" on the same personality test. So the instrument is just not reliable, and yet people are publishing papers based on it.
Jane: Yeah, and that's the core problem the paper is trying to fix. It's not anti-AI. It's pro-rigor. The author wants us to be able to make genuine discoveries about these systems, but we need the right tools to separate real computational phenomena from artifacts.
Tom: So the title really is the thesis. The workflow is validity-guided, meaning every step is designed to build evidence that your measurement is sound before you make any claims. I'm excited to get into the details of that workflow, Jane. What's coming up next?
Jane: Next we're going to look at how the paper actually defines the different research goals, because that's the crucial first step. You don't validate a text classifier the same way you validate a cognitive model, and the paper has a really clear framework for that.
Tom: And that distinction is going to matter a lot for how you read every other paper in this space. Stay with us.
Summary of the Paper: Jane: So Tom, we're back with "A validity-guided workflow for robust large language model research in psychology," and I want to talk about the core structure of the paper, because it's genuinely useful. The author lays out four different ways you can use an LLM in research, and each one has different validation requirements.
Tom: Four ways. Let's hear them.
Jane: First, you can use the LLM as a research tool. That's things like coding text, classifying sentiment, generating stimuli. Second, you can treat it as an evaluation target, where you're asking questions like, does this model show a stable personality pattern? Third, you can use it as a human simulator, trying to replicate survey responses from a population. And fourth, you can treat it as a cognitive model, claiming the model's processes are analogous to human cognition.
Tom: And the key insight is that these aren't just different use cases. They have escalating evidentiary requirements. If you're using a model to code interview transcripts, you need to show it agrees with human coders. That's a functional claim. But if you're claiming the model has theory of mind, you're making a much bigger claim, and you need much stronger evidence.
Jane: Exactly. And the paper has this great table mapping each research goal to the specific validity evidence required. For a research tool, you need reliability and accuracy. For an evaluation target, you need construct validity, internal structure, convergent and discriminant evidence. For a cognitive model, you need causal interventions, like ablation studies, to show the mechanism is actually doing the work.
Tom: So the paper is essentially saying, stop treating every LLM claim like it's the same kind of claim. If you want to say the model is anxious, you need to show that your anxiety measure is reliable, that it has a coherent factor structure, that it predicts anxiety-related behaviors in other contexts. That's a high bar.
Jane: And that's the right bar. The paper gives a concrete example of why this matters. It cites research showing that personality inventories given to LLMs collapse under factor analysis. Instead of getting the Big Five dimensions, you get one giant factor that looks like verbal fluency. So the model is just producing fluent text, not expressing distinct personality traits.
Tom: So the measurement is capturing language ability, not personality. And if you don't do the factor analysis, you'd never know. You'd publish a paper saying, look, GPT-four is conscientious, when really you've just measured how well it can produce coherent sentences.
Jane: Right. And the paper's summary really emphasizes that this workflow is sequential but iterative. You start by defining your goal, then you validate your instrument, then you design your experiment, then you execute, then you analyze, then you report. But if you discover something at a later stage, you can go back and revise earlier decisions.
Tom: That iterative structure is important because it acknowledges that research isn't linear. You might find that your instrument is unreliable and need to go back to the drawing board. The paper is giving you permission to do that instead of just pushing forward with bad measurements.
Jane: And that's the real contribution here. It's not just a checklist. It's a way of thinking about LLM research that puts measurement quality at the center. And that's going to change how we interpret a lot of the existing literature.
Tom: I can already see how this would reshape the field. But I'm curious about the specific improvements the paper suggests. What does it actually tell you to do at each stage? That's our next segment.
Jane: Yes, and I think the practical guidance there is where this paper really shines. Let's get into it.
Improvements Suggested by the Paper: Tom: So Jane, we've established that the paper gives us this framework of four research goals and escalating validity requirements. Now let's talk about what it actually tells us to do. What are the concrete improvements it's suggesting?
Jane: The paper breaks the workflow into six stages, and the first one is defining your research goal. That sounds obvious, but the paper argues that misclassification is rampant. People claim they're testing theory of mind when they're really just using the model as a tool to answer questions. And that category error cascades through everything else.
Tom: So if you mislabel your goal, you'll under-validate. You'll skip the psychometric testing that would have caught the problem. That makes sense. What's the second stage?
Jane: Stage two is developing and validating the computational instrument. And this is where the paper gets really concrete. For research tools, you need to test agreement with human coders or accuracy against gold standards. But for evaluation targets and cognitive models, you need full psychometric validation. That means content validity, reliability, internal structure, response process evidence, convergent and discriminant validity, and consequential evidence.
Tom: That's a lot of evidence. Can you give me an example of what one of those looks like in practice?
Jane: Sure. Take response process evidence. The paper describes a technique where you use chain-of-thought prompting to see how the model arrives at its answer. If a model passes a theory-of-mind test but its reasoning trace shows it's just retrieving a similar example from training data, then your measurement is contaminated. You haven't measured theory of mind. You've measured memorization.
Tom: So you're literally inspecting the model's reasoning to make sure it's engaging with the construct, not just pattern matching. That's a really practical improvement over just looking at final accuracy scores.
Jane: Exactly. And the paper also emphasizes parallel forms reliability. You create semantically equivalent but syntactically different versions of your prompts. If the model's answers change dramatically when you switch from "A, B, C" to "one two three" then your instrument is measuring formatting artifacts, not the construct.
Tom: That's the "on a box" versus "in a box" problem from the intro. The paper is giving you a systematic way to catch that before you publish.
Jane: Right. And then stages three through five are about designing experiments, executing them, and analyzing the data. The paper has really important guidance there too. For example, it warns about the non-independence problem. If you query the same model a hundred times, those responses are not independent observations. They're clustered within one system. Treating them as independent inflates your false positive rate.
Tom: So you need multilevel modeling or cluster-robust standard errors. The paper is bringing standard statistical tools to bear on LLM data, which is exactly what's been missing.
Jane: And stage six is about reporting. The paper recommends pre-registering your instrument validation plan, documenting exact model versions and parameters, and storing raw outputs before any parsing. Because models get deprecated, and if you don't save the raw data, you can never audit your results later.
Tom: That's the transparency piece. And I love that the paper frames validation as an ongoing process, not a one-time achievement. A model that's validated in July might behave differently after a silent update in August.
Jane: Exactly. And that's why the paper suggests building shared repositories of validated instruments. So researchers don't have to start from scratch every time a new model comes out. They can adapt existing validation protocols to new versions.
Tom: That would be a huge efficiency gain for the field. And it would also make it possible to track how model behavior evolves over time, which is scientifically interesting in its own right.
Jane: Right. So the improvements here are really about slowing down, being methodical, and building cumulative knowledge instead of chasing one-off findings. And I think that's the message we should carry into the conclusion.
Conclusion: Tom: So Jane, we've spent this whole episode on "A validity-guided workflow for robust large language model research in psychology," and I think we should wrap up by pulling together what we've learned.
Jane: Yeah, let's do that. The paper's central message is that LLM research in psychology is suffering from a measurement crisis. People are publishing claims about model personality, theory of mind, moral reasoning, and those claims often fall apart when you apply basic psychometric scrutiny.
Tom: And the solution the paper offers is this six-stage workflow that starts with defining your research goal and scales your validation requirements accordingly. If you're building a tool, validate functionally. If you're making claims about psychological properties, do the full construct validation.
Jane: The paper also gives us a great vocabulary for talking about these problems. Measurement phantoms, cognitive phantoms, the alignment-as-explanation fallacy. These are terms that should become standard in the field because they name real phenomena that we keep stumbling over.
Tom: And I think the biggest takeaway for me is that this isn't anti-AI research. It's actually pro-discovery. By validating our instruments properly, we can find the genuine computational phenomena that are hiding behind all the artifacts. The paper's example of "LLM selfhood" shows how you can reconceptualize a human construct for a disembodied system and measure something real and stable.
Jane: Right, and that's the hopeful message. The workflow isn't a barrier to research. It's the path to research that will actually last. Because right now, a lot of what's being published is going to age terribly. Models get updated, prompts get changed, and the findings evaporate. This workflow is about making findings that survive contact with reality.
Tom: And that matters beyond academia. If we're going to use LLMs in clinical settings, in education, in policy, we need to know that the psychological measurements we're taking are sound. Otherwise we're building applications on sand.
Jane: Absolutely. So we're saying goodbye to this paper, but I think its influence is going to be felt for a long time. It's a call for rigor, and it gives you the tools to achieve it.
Tom: Well said, Jane. That's our show for today. Thanks to everyone listening, and we'll be back soon with another paper to break down. Until then, keep questioning your measurements.
Jane: And keep your prompts stable. See you next time.
Zhicheng Lin
Yonsei University · University of Science and Technology of China
cs.HC, cs.AI, cs.CL, cs.CY
Submitted: 2026-08-15
Updated: 2026-08-18
Journal ref: Behavior Research Methods, 58, 216 (2026)
DOI: 10.3758/s13428-026-03073-2
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 69/100
The gist: The paper presents a six-stage workflow for conducting robust psychological research using large language models (LLMs), guided by a "dual-validity framework" that integrates psychometric validation
Key concepts
- Measurement phantoms
- These are statistical artifacts that appear as real psychological phenomena in LLM research but are actually just the model reacting to formatting quirks or patterns in its training data. They look like genuine psychological findings but lack substance.
- Dual-validity framework
- This framework combines psychometric tradition with causal inference standards. It requires testing not only if a scale is statistically sound but also whether changes can be attributed to the manipulation rather than computational confounding factors.
- Research goals and validation requirements
- The paper outlines four ways to use an LLM in research, each needing different validation. For example, using it as a cognitive model requires causal interventions like ablation studies, whereas using it as a research tool needs reliability and accuracy.
- Response process evidence
- This involves inspecting how the model arrives at an answer, often through chain-of-thought prompting. If the reasoning trace shows pattern matching instead of engaging with the construct, the measurement is considered contaminated.
Terminology
Summary
The paper presents a six-stage workflow for conducting robust psychological research using large language models (LLMs), guided by a dual-validity framework
that integrates psychometric validation with causal inference. The workflow is motivated by evidence of severe measurement unreliability in LLM research, where personality assessments collapse under factor analysis, moral preferences reverse with punctuation changes, and theory-of-mind accuracy varies widely with trivial rephrasing.
These are termed measurement phantoms—statistical artifacts masquerading as psychological phenomena.
The workflow's logic is sequential but iterative,
enforcing a measurement-first approach: One cannot design robust experiments (Stage 3) without validated instruments (Stage 2), and one cannot validate instruments without clarity about research ambitions (Stage 1).
Stage 1: Define Research Goal. Researchers must explicitly classify their claim into four categories with escalating validity requirements
: (1) LLMs as research tools
(functional performance claims), (2) LLMs as evaluation targets
(descriptive claims about stable behavioral patterns), (3) LLMs as human simulators
(comparative claims that outputs statistically parallel human populations), and (4) LLMs as cognitive models
(mechanistic claims about structural and causal analogy to human cognition). Misclassification cascades through subsequent stages—studies claiming LLMs have theory of mind while providing only tool-level validation commit this category error.
A table maps each goal to required validity evidence (e.g., content validity, reliability, internal structure, response process evidence, convergent/discriminant validation, consequential evidence, and causal inference validity types).
Stage 2: Develop and Validate the Computational Instrument. Validation diverges into two pathways. Pathway A (Research Tools) requires task-specific validation: for classification/coding tasks, evidence includes agreement with human coders, accuracy against gold-standard datasets, or consistency with established psychometric instruments,
using a two-stage protocol of prompt refinement on a manually coded subset followed by evaluation on a held-out test set. For stimulus generation, evidence centers on content validity, assessed through expert review, user ratings, and—crucially for experimental stimuli—manipulation checks.
For text manipulation, evidence involves assessment of both preservation and transformation,
including semantic fidelity and functional equivalence testing.
Pathway B (Full Psychometric Validation) is required for evaluation targets, human simulators, and cognitive models. It unfolds in three phases. Phase 2a (Content Validity and Instrument Development) requires reconceptualizing human-centric constructs for disembodied systems
—e.g., Conscientiousness... cannot mean effortful self-discipline in an LLM. Instead, it might manifest as systematic response organization, meticulous instruction-following, or structured output formatting.
Researchers must generate an item pool, conduct expert review to mitigate training data contamination,
linguistic confounds,
and instruction contamination,
and perform pilot testing using chain-of-thought prompting for a qualitative check to see if the model is engaging with the construct-relevant aspects.
Phase 2b (Reliability Assessment) establishes measurement stability. Test-retest reliability is reframed as characterizing response variance—the inherent stochasticity in its generation process for a given input,
using API access with exact model versions and controlled parameters, generating score distributions over multiple independent runs
(e.g., 20-100 runs until a prespecified stability criterion). Parallel forms reliability tests semantically equivalent but syntactically distinct prompt variants
to detect hypersensitivity. Internal consistency uses McDonald’s Omega (ω)
over Cronbach’s alpha, while being wary of artificially high consistency
signaling a response set.
Parameter stability is tested across temperature, top p, and top k,
and across model updates for reusable instruments.
Phase 2c (Construct Validity Assessment) builds evidence through multiple sources: internal structure analysis via factor analysis (noting that multi-factor structures collapse into a single, undifferentiated dimension resembling verbal fluency
), response process investigation using chain-of-thought and systematic prompt perturbation
(with examples from theory-of-mind research showing how format changes revealed a 'hyperconservative' response policy
or a simple heuristic—rigidly attributing ignorance regardless of context
), convergent/discriminant validation (noting that LLMs frequently fail this validation through a systematic disconnect: 'Self-reported' scores on explicit questionnaires fail to predict behavioral manifestations in interactive tasks,
but that valid nomological networks remain achievable through methodologically appropriate approaches
like adapted Implicit Association Tests), and consequential evidence analysis, which includes behavioral predictivity
and bias and representation audit,
stakeholder consultation,
and define appropriate use boundaries.
Stage 3: Design Experiment. With a validated instrument, researchers must Operationalize the Manipulation and Outcome,
ensuring the manipulation must instantiate the theoretical cause, not merely its linguistic correlates.
They must control four categories of validity threats: internal validity (addressing prompt-level confounds
like positional effects and scenario reconstruction
where LLMs actively infer entire contexts from minimal cues,
plus technical confounds
like silent model updates and context accumulation), external validity (defining boundaries across models, human populations, and time, noting temporal displacement
), construct validity of the manipulation (guarding against the alignment-as-explanation fallacy
using factorial designs that cross substantive manipulations with irrelevant variations), and statistical conclusion validity (addressing non-independence
and the 'correct answer' effect
that reduces response variability). Finally, researchers must Develop Pre-registration Plan
to counter HARKing and p-hacking,
using a multi-stage registration process that first pre-registers instrument development, then updates with the final experimental protocol.
Stage 4: Execute and Document the Experiment. This stage requires Specify and Document the Environment,
including the specific model endpoint (e.g., gpt-4-0125-preview), technical parameters (temperature, top p, max tokens), the date and time of data collection,
and for web interfaces, documenting the model family selected... interaction dates and times, and capture screenshots of interface settings.
Execution must be with Transparency,
allowing deviations from pre-registration provided they are explicitly documented and justified,
with the original analysis still reported. Ensure Data Preservation
requires storing the full prompts exactly as sent to the system; the complete, raw model outputs before any parsing; all relevant metadata, such as timestamps and token counts; and any code used for pre- or post-processing,
deposited in a public repository.
Stage 5: Analyze and Interpret Results. Researchers must Perform Data Quality and Assumption Checks
(screening for anomalies and checking statistical assumptions). They must Address Data Non-Independence
through three strategies: aggregation
(collapsing responses into summary statistics), multilevel modeling
or generalized estimating equations (GEE),
and cluster-level bootstrap
that resamples entire models. For grouped prompts, use cluster-robust standard errors,
and for model comparisons, paired analyses of question-by-question score differences.
Always report the number of clusters and intracluster correlation coefficients.
They must Conduct Robustness Analyses,
including both technical robustness (parameter settings, prompt formatting, data parsing) and conceptual robustness
(testing against theoretically relevant variations in task logic and content
). Finally, Calibrate Interpretations of Effects,
focusing on the consistency of patterns across robustness checks rather than the magnitude or p-value of any single analysis.
Stage 6: Report and Reconceptualize. Reporting must be Transparent and Accessible,
following guidelines like TRIPOD-LLM
or MI-CLEAR-LLM,
with a replication package including Model and Environment Specifications,
Handling of Stochasticity,
and Prompt and Data Documentation
(including a statement on potential data contamination
). Researchers must Constrain Claims to Evidence,
avoiding anthropomorphic language and distinguish between observed behavioral performance and inferred underlying competence.
They should Use Findings to Reconceptualize and Refine
theories and methods, and Address Limitations and Ethical Implications,
including generalizability across model versions and broader societal impacts.
The paper illustrates the workflow with an example of LLM selfhood,
reconceptualizing the construct as the stability and coherence of self-referential linguistic patterns across different contexts and tasks,
and validating it through reliability, construct validity, and response process checks before designing a priming experiment.
The concluding remarks emphasize that the workflow is designed not as a restrictive checklist, but as a generative framework,
that it forces them to abandon the naive anthropomorphic label and build a new, computationally grounded construct from the ground up,
and that it deliberately introduces friction into the research process, slowing down the rapid cycle of 'prompt-and-publish' in favor of a more methodical, front-loaded validation process.
The paper calls for shared repositories
of validated computational instruments to serve as methodological blueprints,
enable historical benchmarking,
and preserve conceptual work,
concluding that Methodological rigor is not an obstacle to discovery but the only path toward it.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved system can now do:
-
Implementation: Add a validation layer that tests response stability across semantically equivalent prompt variants (e.g., changing option labels from
A/B
to1/2
, altering punctuation, reordering options) before finalizing outputs. -
Capability: The system can now flag when its own responses are brittle—i.e., when trivial formatting changes would produce materially different answers. It can then either request clarification, default to a more robust response, or explicitly warn the user that the result is prompt-sensitive.
-
Implementation: Integrate a psychometric self-check module that, before claiming a psychological trait (e.g.,
I am conscientious
), evaluates internal consistency across multiple self-referential probes (e.g., behavioral scenarios, forced-choice items, self-descriptions) and checks for contradictions (e.g., endorsing bothI am an introvert
andI am an extrovert
). -
Capability: The system can now distinguish between stable, coherent behavioral patterns and statistical artifacts. It will refuse to make trait claims when evidence is inconsistent, instead reporting
no stable pattern detected
or offering a probabilistic statement with confidence intervals. -
Implementation: Embed a confound-checking routine that, when given a task (e.g.,
simulate a moral dilemma
), automatically generates factorial variations (e.g., changingCase 1
to(A)
, altering scenario details) and tests whether the core output changes. If the output shifts significantly, the system flags the task as confounded and suggests a more robust prompt design. -
Capability: The system can now proactively identify when a user's prompt is likely to produce artifacts rather than genuine reasoning, and can recommend prompt modifications (e.g., adding explicit constraints, randomizing option order) to improve validity.
-
Implementation: When generating multiple responses to the same prompt (e.g., for simulation or survey replication), the system will automatically apply cluster-level corrections (e.g., cluster-robust standard errors, multilevel modeling) and report intracluster correlation coefficients. It will also warn users when treating repeated responses as independent observations would inflate false-positive rates.
-
Capability: The system can now produce statistically defensible aggregate outputs (e.g., simulated survey results) that account for within-model dependency, reducing the risk of spurious findings in downstream analyses.
-
Implementation: Maintain an internal log of model version, training data cutoff, and API parameters (e.g., temperature, top p) for every response. When a user queries the same prompt across different sessions, the system will compare outputs and report drift (e.g.,
Your results may differ from a previous run on 2025-07-01 due to model update
). -
Capability: The system can now warn users about temporal instability, helping them avoid overgeneralizing findings from a single snapshot. It can also suggest re-validation if the model version has changed since the last reliable measurement.
-
Implementation: Automatically generate a structured replication package for any research-oriented task, including: exact prompts, raw outputs (pre-parsing), model version, date/time, parameter settings, and a statement on potential training-data contamination (e.g.,
This prompt may resemble publicly available theory-of-mind tests
). -
Capability: The system can now produce audit-ready outputs that comply with emerging reporting guidelines (e.g., TRIPOD-LLM, MI-CLEAR-LLM), enabling other researchers to replicate or re-analyze results without ambiguity.
-
Implementation: When a user asks the system to measure a human-centric construct (e.g.,
personality
) in an LLM, the system will offer a reconceptualized, computationally grounded alternative (e.g.,systematic response organization
instead ofconscientiousness
) and explain the ontological mismatch. -
Capability: The system can now guide researchers toward more valid measurement targets, reducing the risk of
measurement phantoms
and improving the interpretability of results. -
Self-validate: Detect and flag its own prompt sensitivity and internal contradictions before presenting results.
-
Statistically correct: Automatically adjust for non-independence in repeated responses, reducing false positives.
-
Causally cautious: Identify confounded prompts and suggest more robust experimental designs.
-
Temporally aware: Warn users about model version drift and data cutoff limitations.
-
Fully auditable: Generate complete replication packages with all metadata, enabling transparent and reproducible research.
-
Conceptually precise: Avoid anthropomorphic overclaims by offering computationally grounded alternatives when human constructs do not apply.
These improvements directly operationalize the paper's six-stage workflow, making the AI system not just a tool for generating text, but a reliable, validity-aware research instrument.
Abstract
Large language models (LLMs) are rapidly being integrated into psychological research as research tools, evaluation targets, human simulators, and cognitive models. However, recent evidence reveals severe measurement unreliability: Personality assessments collapse under factor analysis, moral preferences reverse with punctuation changes, and theory-of-mind accuracy varies widely with trivial rephrasing. These "measurement phantoms"--statistical artifacts masquerading as psychological phenomena--threaten the validity of a growing body of research. Guided by the dual-validity framework that integrates psychometrics with causal inference, we present a six-stage workflow that scales validity requirements to research ambition--using LLMs to code text requires basic reliability and accuracy, while claims about psychological properties demand comprehensive construct validation. Researchers must (1) explicitly define their research goal and corresponding validity requirements, (2) develop and validate computational instruments through psychometric testing, (3) design experiments that control for computational confounds, (4) execute protocols with transparency, (5) analyze data using methods appropriate for non-independent observations, and (6) report findings within demonstrated boundaries and use results to refine theory. We illustrate the workflow through an example of model evaluation--"LLM selfhood"--showing how systematic validation can distinguish genuine computational phenomena from measurement artifacts. By establishing validated computational instruments and transparent practices, this workflow provides a path toward building a robust empirical foundation for AI psychology research.
Sources
- A Survey on Data Contamination for Large Language Models
- Does It Make Sense to Speak of Introspection in Large Language Models?
- The Capability of Large Language Models to Measure Psychiatric Functioning
- Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina
- The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective
- LLM-based Text Simplification and its Effect on User Comprehension and Cognitive Load
- Thinking beyond the anthropomorphic paradigm benefits LLM research
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations
- Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
- From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology
- Large Language Models as Psychological Simulators: A Methodological Guide
- Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- Multi-ToM: Evaluating Multilingual Theory of Mind Capabilities in Large Language Models
- Personality Traits in Large Language Models
- From traces to measures: Large language models as a tool for psychological measurement from text
- Challenging the Validity of Personality Tests for Large Language Models
- Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
- Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
- Evaluating General-Purpose AI with Psychometrics
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support