InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation".
Jane: Simulating real personalities with large language models requires grounding generation in authentic personal data,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the title and who put this together: "InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation." The name really tells you exactly what they're trying to accomplish—creating a scalable way to make AI personalities that sound like real people through interview data.
Jane: And it’s impressive seeing the authors listed, including Yu Li, Pranav Narayanan, Venkit Yada, Pruksachatkun Chien-Sheng Wu from Salesforce Research. Having researchers from a company known for large language models involved suggests this framework is built with the current state of AI in mind.
Lu: That background gives you confidence that the technical foundation will be solid because they are coming from a place deeply involved in building these kinds of systems. It shows they understand the nuances of what makes an interview transcript valuable data for simulation, not just any random text.
Meng: I’m more interested in the "scalable framework" part. Scaling up to one thousand personalities and six hundred seventy-one thousand question-answer pairs is a massive undertaking. How do you think they managed the infrastructure needed to process that much structured data efficiently?
Lalam: That scalability comes from their rigorous data curation pipeline described in the paper. They didn't just dump transcripts in; they built a reproducible system for selecting, verifying, and splitting that data into training and test sets, which is crucial for reliable evaluation.
Tom: So it’s not just about having a big dataset; it’s about having a systematic way to organize that data so you can actually run meaningful tests on the simulation fidelity. It moves the goal from "can the AI sound like this person?" to "how close is this AI's response to what this person actually said in an interview?"
Jane: That shift in focus is exactly what makes it significant; they are grounding the simulation in authentic dialogue rather than just broad biographical facts, which should lead to much more nuanced and realistic interactions.
The paper's summary: Tom: To summarize what "InterviewSim" actually does, it focuses on extracting over six hundred seventy-one thousand question-answer pairs from twenty-three thousand verified interview transcripts across a thousand public personalities. It essentially builds a massive library of real human conversational data specifically designed to train and test personality simulation models.
Jane: That’s a lot of raw material, and the core proposal is this multi-dimensional evaluation protocol with four complementary metrics: content similarity, factual consistency, personality alignment using the Big Five model, and factual knowledge retention. It gives researchers multiple angles to judge how good an AI simulation actually is.
Lu: The summary emphasizes that this framework is method-agnostic; it doesn't matter if you use a simple prompt or a complex retrieval augmentation method; the evaluation protocol remains the same for all of them, which makes it very powerful for comparing different generation techniques.
Meng: I see how important that is for practical application. If we want to deploy an AI assistant that needs to sound like an expert in a specific field, knowing which generation method best preserves factual consistency versus stylistic flair is key information here.
Lalam: The paper highlights the finding that methods grounded in real interview data substantially outperform those relying only on biographical profiles or the model’s parametric knowledge when simulating personalities. That direct comparison against simpler methods is a big selling point for its utility.
Tom: So, it boils down to using real interview data to measure how well an AI captures not just *what* a person knows, but *how* they talk and reason when answering questions in context with their actual expertise.
The paper's improvements: Tom: The authors suggest several improvements to the general approach of simulating these personalities. They propose moving away from relying only on simple biographical profiles toward using long-form conversational content as the primary grounding for simulation, which is a big step in fidelity.
Jane: Another key suggestion they make is that they should focus on structuring the interview data carefully into Q andA pairs across four thematic categories to ensure comprehensive coverage of different aspects of a personality's knowledge and interaction style.
Lu: They also detail how they split the interview data temporally into training and test sets, which is important for ensuring the model can actually generalize its responses rather than just memorizing specific answers from the training set.
Meng: From a practical deployment view, I’m interested in their suggestion about how different generation methods trade off; it points toward making informed choices about whether you need style capture or factual accuracy depending on what you're trying to achieve with the AI.
Lalam: One of the improvements they highlight is the way they designed their multiple choice question metric, which converts complex Q andA pairs into structured questions with specific distractors like negation and misconception, making it a much stricter test than just checking for simple similarity.
Tom: So, beyond just having the data, they are proposing a smarter way to structure that data and a more detailed way to grade the simulation using these four metrics to get a clearer picture of success.
Conclusion: Tom: Wrapping things up with "InterviewSim," the paper demonstrates that grounding personality simulation in real interview data yields substantial improvements over methods based only on profiles or parametric knowledge, and this is achieved through their multi-dimensional evaluation protocol.
Jane: It really confirms that for high-fidelity simulation, authentic conversational transcripts are a much stronger foundation than anything else currently available for this task. This moves the field closer to creating AI that interacts with humans in a way that feels genuinely representative of another person's communication style.
Lu: The implication here is that we can build models capable of capturing complex, nuanced ways people think and communicate when they are actually being interviewed, which opens up possibilities for much more sophisticated interactive systems down the road.
Meng: I think the practical takeaway is that researchers need to be very deliberate about their simulation goals—if you prioritize factual consistency in a high-stakes area, you need to lean into the chronological grounding methods they tested.
Lalam: To finish up, "InterviewSim" provides a clear blueprint for building and evaluating these agents systematically, offering actionable insights for anyone looking to move beyond superficial personality modeling and toward truly grounded simulation.
Salesforce Research
cs.CL, cs.AI, cs.CY
Submitted: 2026-02-23
Updated: 2026-10-01
Comments: Accepted to COLM 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 91/100
The gist: Simulating real personalities with large language models requires grounding generation in authentic personal data, and this framework addresses the gap by introducing an interview-grounded evaluation
Key concepts
- InterviewSim Framework
- This is the entire system designed to simulate personalities. It involves two parts: a pipeline to gather and clean massive datasets from interviews, and an evaluation protocol with four specific metrics to check how well an AI simulation matches the real person's interview data.
- Content Similarity
- This metric measures how closely a generated response matches the actual information and ideas found in the real interview. A high score means the AI captured the core content accurately, scoring on a scale of 1 to 5 where 5 is perfect similarity.
- Factual Consistency (Contradiction Ratio)
- This assesses whether an AI's response contradicts established facts about the personality. It classifies answers as Entailment, Neutral, or Contradiction based on interview summaries and calculates a contradiction ratio to measure factual accuracy.
- Chronological-based Generation
- This generation method uses in-context learning by feeding the model actual interview Q&A pairs in sequence. This approach is found to be superior for preserving the temporal and contextual structure of the original interviews, leading to better factual consistency.
Terminology
Summary
Simulating real personalities with large language models requires grounding generation in authentic personal data, and this framework addresses the gap by introducing an interview-grounded evaluation protocol for personality simulation at a large scale. The core finding is that methods grounded in real interview data substantially outperform those relying solely on biographical profiles or the model’s parametric knowledge.
The gist
We extract over 671,000 question-answer pairs from 23,000 verified interview transcripts across 1,000 public personalities and propose a multi-dimensional evaluation framework with four complementary metrics measuring content similarity, factual consistency, personality alignment, and factual knowledge retention.
The INTERVIEWSIM Framework
The framework consists of two main components: (1) a reproducible pipeline for curating large-scale interview datasets with structured train/test splits, and (2) a multi-dimensional evaluation protocol that assesses simulation fidelity against held-out interview responses. The data collection process involves selecting 1,000 public personalities across eight occupational categories and curating targeted corpora of longform conversational content. Quality control is rigorous, employing an automated stage with GPT-4.1 followed by human verification by a team of trained annotators to ensure accuracy, achieving a retention rate of 73.5% for verified transcripts.
Evaluation Protocol
The framework utilizes four complementary metrics to assess simulation fidelity:
-
Content Similarity: Measures how well a generated response captures the same information and ideas as the ground truth answer from the personality’s actual interview, scored on a 1-5 scale where 5 indicates high similarity.
-
Factual Consistency: Evaluates whether generated responses contradict established facts about the personality by classifying them as Entailment, Neutral, or Contradiction based on a fact summary of the interviews. The metric computed is the contradiction ratio (CR).
-
Personality Similarity: Assesses trait alignment using the Big Five (OCEAN) model, comparing a reference profile derived from held-out test set answers against the generated profile using ordinal distance.
-
Multiple Choice Question (MCQ): A knowledge-based metric that converts complex Q&A pairs into atomic questions with structured distractors (opposite/negation, near-miss, misconception), and measures performance via accuracy and a reward metric that penalizes confident factual errors more heavily than near-miss mistakes.
Generation Methods Comparison
The study systematically compares four generation methods: Simple Prompt (relying on parametric knowledge), Wiki-based (augmenting with biographical profiles), Chronological-based (using in-context learning with actual interview Q&A pairs), and Memory-based retrieval (a retrieval-augmented method using top-k most similar training examples). Empirical findings reveal a clear trade-off: retrieval-augmented methods excel at capturing personality style and response quality, while chronological-based methods better preserve factual consistency and knowledge retention.
Specifically, memory-based retrieval achieved the highest content similarity of 3.50 and personality similarity of 78.4%, while chronological-based methods achieved the lowest contradiction ratios, decreasing from 6.18% at 100 examples to 5.70% at 500 examples.
Key Empirical Insights
Analysis across question types and personality categories reveals systematic variations in difficulty: Social Identity questions yield the highest contradiction rates, requiring precise numerical facts, while Motivations and Values questions show the lowest contradiction rates due to the flexibility of phrasing. Similarly, Science & Academia personalities achieve the lowest contradiction rates. The trade-off between retrieval and sequence is further highlighted: Memory-based retrieval prioritizes topical relevance,
capturing stylistic features like emotional intensity, but chronological-based prompts, by contrast, preserve the temporal and contextual structure of original interviews.
Furthermore, MCQ robustness analyses show that chronological interview grounding consistently achieves higher or comparable accuracy and reward across all data regimes.
Limitations and Scope
The study acknowledges limitations regarding dataset scope (primarily Western media from 2015 to 2024), potential biases in LLM judges, and the fact that the held-out test set is drawn from the same time period as training data, preventing assessment of personality evolution over time. Methodologically, while retrieval-augmented methods excel at style capture, their performance on factual consistency is sensitive to context depth. The research concludes by establishing a foundation for principled method selection based on application requirements and provides actionable insights into advancing personality simulation research.
Ethics and Data Curation
The study adheres to strict privacy protocols by using only publicly available interview records, omitting specific platform details and mapping subject names to anonymized IDs. The complete curation pipeline and evaluation protocol are released so researchers can independently reproduce the dataset, mitigating risks associated with impersonation or unauthorized commercial exploitation. Human annotation was critical for quality assurance, employing a risk-based QA strategy to ensure reliability across the large corpus.
Improvements for AI systems
Here are specific, actionable improvements for AI systems based on the INTERVIEWSIM framework:
) Improved System Capabilities & Applications:
The core improvement is shifting from generic, demographically-grounded simulation to high-fidelity, real-world behavior replication grounded in authentic human dialogue. The resulting improved AI system can perform the following:
Sources
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Simulating Social Media Using Large Language Models to Evaluate Alternative News Feed Algorithms
- Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History
- The Need for a Socially-Grounded Persona Framework for User Simulation
- CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
- OASIS: Open Agent Social Interaction Simulations with One Million Agents
- SocioVerse: A World Model for Social Simulation Powered by LLM Agents and A Pool of 10 Million Real-World Users
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering