Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model
cs.CL
Submitted: 2026-08-19
Updated: 2026-09-13
Comments: v2: restructured into standard paper format (Method / Results I-II / Conclusion / lettered appendix); no changes to numbers or claims
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differences, and whether this reflects what a model knows or
Terminology
Abstract
Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differences, and whether this reflects what a model knows or uses has remained untested. Using representational similarity analysis against Pew American Trends Panel ground truth, we score demographic read-out locations in Mistral-7B and intervene causally across six attribute types. The internal geometry is faithful: attention-head read-outs dominate the standard residual read-out, reaching selection-corrected ρ up to 0.63 -- about 70% of the measurement-reliability ceiling -- and one head, L11 H16, is significantly faithful across all six types, though race-based types stay weak and prompt-fragile, replicating in a second model family. Yet causal use does not track fidelity: the clearest causal pathway (p=0.002) sits in one of the least faithful types, the most faithful type shows no correction-surviving effect, and full identity swaps in the prompt move predictions by under 2% of their error. A 128-dimensional probe on that head lands 21-31% closer to survey truth than the model's answers, yet recovers almost none of the per-question group ordering. Readable, faithfully arranged, and causally used are three dissociable properties of the same model; treating them as one claim is what keeps the "can LLMs simulate populations" debate unresolved.
Sources
- How Well Do Large Language Models Capture Human Personality?
- Linear socio-demographic representations emerge in Large Language Models from indirect cues
- How Reliable are Causal Probing Interventions?
- Assessing Cross-Cultural Alignment between ChatGPT and Human Societies: An Empirical Study
- Localizing Persona Representations in LLMs
- Questioning the Survey Responses of Large Language Models
- Towards Measuring the Representation of Subjective Global Opinions in Language Models
- Social Chemistry 101: Learning to Reason about Social and Moral Norms
- Probing Persona-Dependent Preferences in Language Models
- How to use and interpret activation patching
- Quantifying the Persona Effect in LLM Simulations
- Mistral 7B
- Linear Representations of Political Perspective Emerge in Large Language Models
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
- Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics
- Locating and Editing Factual Associations in GPT
- Dissecting Persona-Driven Reasoning in Language Models via Activation Patching
- Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
- Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection
- From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering