Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

arXiv:2608.18768 · cs.CL · Submitted 2026-08-19 · Read on arXiv

cs.CL

Submitted: 2026-08-19

Updated: 2026-09-13

Comments: v2: restructured into standard paper format (Method / Results I-II / Conclusion / lettered appendix); no changes to numbers or claims

License: http://creativecommons.org/licenses/by/4.0/

The gist: Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differences, and whether this reflects what a model knows or

Terminology

Abstract

Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differences, and whether this reflects what a model knows or uses has remained untested. Using representational similarity analysis against Pew American Trends Panel ground truth, we score demographic read-out locations in Mistral-7B and intervene causally across six attribute types. The internal geometry is faithful: attention-head read-outs dominate the standard residual read-out, reaching selection-corrected ρ up to 0.63 -- about 70% of the measurement-reliability ceiling -- and one head, L11 H16, is significantly faithful across all six types, though race-based types stay weak and prompt-fragile, replicating in a second model family. Yet causal use does not track fidelity: the clearest causal pathway (p=0.002) sits in one of the least faithful types, the most faithful type shows no correction-surviving effect, and full identity swaps in the prompt move predictions by under 2% of their error. A 128-dimensional probe on that head lands 21-31% closer to survey truth than the model's answers, yet recovers almost none of the per-question group ordering. Readable, faithfully arranged, and causally used are three dissociable properties of the same model; treating them as one claim is what keeps the "can LLMs simulate populations" debate unresolved.

Sources

Related papers