Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
cs.AI, cs.CL, cs.CY
Submitted: 2026-08-23
Updated: 2026-08-23
Comments: 20 pages, 4 figures, 6 tables. Code and derived data available at https://github.com/heathriel/synthetic-memoir-audit
Code: https://github.com/heathriel/synthetic-memoir-audit
License: http://creativecommons.org/licenses/by/4.0/
The gist: When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of
Terminology
Abstract
When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search. The subject and the author of this paper are the same person: a 366-day "page-a-day" book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day's quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis. We define the verification-failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4-98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift - real people, employers, and settings inside invented scenes - though its measured share varies across raters. Independent re-rating replicates the headline (no evidence the original rate was inflated) while showing that the four-way taxonomy has only fair-to-moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject's corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect.
Sources
- Generative Agents: Interactive Simulacra of Human Behavior
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Beyond Profile: From Surface-Level Facts to Deep Persona Simulation in LLMs
- Conversational AI Powered by Large Language Models Amplifies False Memories in Witness Interviews
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
- Large Language Models Hallucination: A Comprehensive Survey
- Confabulation: The Surprising Value of Large Language Model Hallucinations
- ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- Long-form factuality in large language models
- PsyDT: Using LLMs to Construct the Digital Twin of Psychological Counselor with Personalized Counseling Style for Psychological Counseling
- Who's Asking? Evaluating LLM Robustness to Inquiry Personas in Factual Question Answering
- RoleBreak: Character Hallucination as a Jailbreak Attack in Role-Playing Systems
- SHARP: Unlocking Interactive Hallucination via Stance Transfer in Role-Playing LLMs
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection