Evaluating the Efficacy of LLMs to Emulate Realistic Human Personalities

arXiv:2402.14879 · cs.CL, cs.AI · Submitted 2026-08-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Evaluating the Efficacy of LLMs to Emulate Realistic Human Personalities".

Jane: The paper was written by W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang et al. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: So, in "Driving Generative Agents With Their Personality," the team summarizes a lot of groundwork laid by existing Affective Computing systems. They show that it’s not just about emotion anymore, but it’s about the depth and structure behind those emotions.

Jane: The authors used established frameworks like the Big Five personality model—which we all know involves traits such as openness and conscientiousness—to create a blueprint for what a character's personality should be.

Lu: And instead of just stating that "this character is nice," they are defining that person by their specific scores across these five factors, which is a much finer resolution than previous methods allowed.

Meng: The summary highlights the use of the International Personality Item Pool, or IPIP, which provides this massive dataset to ground those abstract personality concepts in real human behavior.

Lalam: This is where it shows that we are moving beyond just emotional reactions; we are building a foundational psychological profile for how the AI thinks and responds to our input.

Tom: It’s impressive how they’ve managed to take this raw, real-world data and turn it into a usable format for generative AI.

Jane: It shows that you can synthesize complex human traits into something manageable without losing the essential characteristics of the original personality type.

Lu: This process allows us to simulate personality in a way that mimics how humans actually categorize and understand individual differences.

Meng: By using IPIP, we’ are ensuring that these digital personas are based on empirical data, which is necessary for practical implementation in games.

Lalam: This is about giving the AI a reliable internal logic, making its responses feel less like random text and more like genuine character expression.

Improvements: Tom: The authors argue that by providing an LLM with this detailed personality profile, we can dramatically improve the quality of NPC interactions compared to what was achievable before these methods were standardized.

Jane: The real improvement lies in consistency; the character won’t suddenly act completely out of character just because a prompt changed, which has been a huge headache for game developers trying to maintain world-building integrity.

Lu: They are providing a robust psychological scaffolding for the LLM, ensuring that its decisions and dialogue reflect its core disposition consistently, which creates incredibly reliable behavioral patterns.

Meng: The practical improvement is in the reliability of the output; we aren't just getting random text anymore, we’re getting content generated by a character whose mental architecture is already established.

Lalam: This allows us to see characters who can evolve in a meaningful way, adapting their personality traits as they interact with the player, which makes the entire experience feel much more organic and alive.

Tom: So, this isn' not just about better dialogue; it’s about creating a psychological consistency that elevates the entire gameplay experience.

Jane: It’s great to hear people talk about developer headaches because that is exactly what this solves—a consistent character is a solvable problem for AI design.

Lu: The LLM becomes a true extension of the character's psyche, not just a tool responding to keywords, which is the ultimate goal of advanced AI integration.

Meng: If we can trust the consistency of an NPC’s actions, we can build much more complex game systems around them without worrying about sudden behavioral shifts.

Lalam: We are moving toward characters that feel like they have their own internal logic and that every step they take is consistent with who they are.

Results: Tom: The results section of "Driving Generative Agents With Their Personality" shows some truly impressive data points regarding how well these models can actually execute a given personality.

Jane: They tested several leading LLM models, and it seems like GPT-four-six hundred thirteen really stood out, achieving an accuracy rate that was incredibly high in embodying the assigned profile.

Lu: That leap from the earlier models is a testament to how far LLMs have advanced; it proves that current generation AI can handle nuanced psychological mapping far better than previous ones.

Meng: The use of Root Mean Squared Prediction Error, or RMSPE, gave us concrete proof that gpt-four-six hundred thirteen is generating responses statistically indistinguishable from the real data profiles.

Lalam: This performance validates the idea that AI can truly internalize a personality, moving beyond just being a mimic to becoming an accurate reflection of complex human characteristics.

Tom: It’s not just a theoretical improvement; we have quantitative proof that this is working exceptionally well across different models and approaches.

Jane: The data really shows us where the current state-of-the-art is, which is helpful for developers planning their next steps in AI implementation.

Lu: This accuracy confirms that the Big Five framework provides a powerful, measurable structure for how AI interprets complex human personality traits.

Meng: For me, it means we can now have high confidence in the behavioral output of an NPC, knowing the statistical error is minimal compared to other models.

Lalam: We can finally see characters whose internal life matches their real-world psychological counterparts, which is a huge step for cultural realism.

Conclusion: Tom: As we wrap up "Driving Generative Agents With Their Personality," it’s clear that integrating psychometric data with LLMs opens up incredible possibilities for creating highly realistic NPCs.

Jane: It feels like we're on the cusp of creating characters that will not only have believable personalities but also possess depth and emotional complexity, something I hope to see in future games.

Lu: My biggest takeaway is the potential for dynamic storytelling, where the AI doesn't just react to events but reacts based on its internal psychological constraints, leading to complex narratives.

Meng: We need to focus on how this could scale into a practical implementation—how we manage that robust dataset and integrate it smoothly into a game engine architecture.

Lalam: It's fundamentally about improving the human experience; making our interactions with digital characters feel less like programming and more like genuine conversation, which is a huge cultural leap.

Tom: I think this research just opens the door to an era where AI can perfectly reflect human personality traits in games, creating truly engaging worlds.

Lu: I'm excited to see how we use these concepts for different psychological markers beyond the Big Five in future iterations of this work.

Meng: The engineering challenge is exciting because scaling a large-scale personality dataset is a manageable, albeit complex, task.

Lalam: It allows us to design characters that feel like they have their own internal lives and that interact with others on equal footing.

Tom: Thank you all for this incredible discussion about "Driving Generative Agents With Their Personality." We're really looking forward to seeing how these innovations will shape the next generation of game development.

W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, Y. Du, C. Yang, Y. Chen, Z. Chen, J. Jiang

cs.CL, cs.AI

Submitted: 2026-08-23

Updated: 2026-08-25

Comments: 11 pages, 4 figures, 3 tables. Published in Proceedings of AIIDE 2024

Journal ref: Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment 20(1) (2024) 65-75

DOI: 10.1609/aiide.v20i1.31867

Project page: https://stephbuon.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 79/100

The gist: This summary details the methodology and findings of the research titled "Driving Generative Agents With Their Personality," which investigates the potential of Large Language Models (LLMs) to

Key concepts

Big Five Personality Model
This framework defines personality using specific traits such as openness and conscientiousness. Instead of simple labels, it provides a blueprint by using precise scores across these five factors, allowing for a much finer resolution when creating complex character profiles for AI design.
IPIP (International Personality Item Pool)
IPIP is a massive dataset used to ground abstract personality concepts in real human behavior. By utilizing this empirical data, researchers ensure that the digital personas are based on evidence, making them suitable for practical implementation in games and ensuring high accuracy.
LLM Emulation Accuracy
The research tests leading LLMs to see how accurately they can embody a given personality profile. Results show that models like GPT-4-6013 achieved high accuracy, generating responses statistically indistinguishable from real data profiles, proving the AI can internalize complex characteristics.

Terminology

Summary

This summary details the methodology and findings of the research titled Driving Generative Agents With Their Personality, which investigates the potential of Large Language Models (LLMs) to utilize psychometric values, specifically personality information, within video game character development.

Introduction and Motivation

The study is motivated by the goal to create NPCs interacting with their environment and displaying a rich tapestry of character depth and emotional complexity. While Affective Computing (AC) systems are already being used by major video game companies to empower Non-Player Characters (NPC) to perceive and project emotions, the current state of LLMs often fails to meet the ambition of generating content that aligns with a specific personality. The research aims to bridge this existing gap by exploring the potential of integrating AC psychometrics with LLMs, thereby fostering a more immersive and emotionally engaging gaming experience.

Methodology: Personality Representation

The core of this inquiry focuses on personality traits, specifically utilizing the Big Five model (also known as the Five Factor or OCEAN model). This model comprises five factors: "Openness to novel experiences, Conscientiousness in tasks and interpersonal relationships, Extraversion in social contexts, Agreeableness towards diverse viewpoints and mutual understandings, and Neuroticism in interpreting circumstances."

The research utilized a dataset derived from the Open-Source Psychometrics Big Five Project (n = 1,015,342 participants). A significant portion of this data—a total of 596,956 test results—was selected for analysis after a thorough data-cleaning process.

The evaluation involved several analytical techniques:

  1. Label Assignment: A nearest-neighbor approach was used to calculate the Euclidean Distance between each LLM test result and the 20 distinct personality profiles, identifying the profile with the smallest distance as the corresponding label.

  2. Dimensionality Measurement: The study projected personalities onto a two-dimensional plane using Cognitive Stability (alpha = A+C+(1-N)) and Cognitive Flexibility (beta = E+O).

  3. Performance Metrics: Two primary metrics were used to assess LLM efficacy:

  • Root Mean Squared Prediction Error (RMSPE): This measures the error in the evaluated test results compared to the baseline dataset.

  • Inter-rater Reliability (IRR): This assesses the extent to which different ratings of the same entity agree with each other.

Results and Analysis of LLM Performance

The study tested a suite of OpenAI models: text-davinci-003, gpt-3.5-turbo-0613, and gpt-4-0613.

The results demonstrated a significant leap in performance among the models:

  • Accuracy: Table 2 shows that while text-davinci-003 achieved an accuracy of 13.8281%, and gpt-3.5-turbo-0613 achieved 17.7734%, gpt-4-0613 exhibited a substantial leap in performance, achieving an accuracy of 73.9844%.

  • Clustering: Figure 2 visually confirmed this, showing that the evaluated test results for gpt-4-0613 move closer to their respective baseline clusters and cluster around its corresponding profile.

  • Error and Consistency: Table 3 (RMSPE) and Table 4 (IRR) further demonstrated that gpt-4-0613 consistently outperformed its predecessors, exhibiting the lowest error rates across various personality profiles, confirming its superior ability to generate responses in line with given personality profiles.

The study concluded that the performance of gpt-4-0613 provided a clear, visual representation of the distribution and underscored its superior capacity to generate responses befitting a given personality profile.

Use Cases for Generative Agents

The research outlines three specific scenarios where an LLM, guided by its assigned personality, can enhance gameplay:

  1. Retelling Stories: The LLM can be prompted to retell a story based on experiences, such as in Tales of Arabian Nights or This War of Mine, with the NPC sharing narratives and interpreting events.

  2. Improvisation: The model can generate content using a yes and... approach, similar to games like Dungeons and Dragons, allowing the NPC to describe actions that align with its psyche.

  3. Narrative Conclusion: The LLM can be tasked with explaining how a story reached its conclusion, using the assigned personality as a heuristic to narrow down choices and fill in narrative gaps, as seen in games like Sherlock Holmes Consulting Detective.

Conclusion

The research concludes that the integration of AC systems and LLMs opens up exciting possibilities for integrating AC systems and LLMs for NPCs, enabling dynamic character transformations. The potential of LLMs to generate content accurately reflecting a provided personality ensures that the resulting NPCs can be created with consistent and believable behavioral patterns, significantly enhancing the overall gaming experience.

Improvements for AI systems

The following improvements are derived from the methodology presented in the paper and represent a robust framework for integrating quantifiable psychological data into Large Language Models, transforming them from simple text generators into consistent, psychologically grounded agents.

1. Structured Psychometric Grounding Layer (AC Integration)

  • Improvement: Instead of relying on vague textual descriptions (Act like a cautious NPC), the system incorporates an external Affective Computing (AC) module that provides a precise, standardized input vector—the 5-tuple (O, C, E, A, N) —which is derived from validated psychometric frameworks (e.g., the Big Five model).

  • Specific Action: This AC layer acts as a non-negotiable constraint on the LLM' behavior. The LLM is explicitly prompted with this numerical data (e g, You are an NPC where Openness is 0.2 and Agreeableness is 0.95 ).

  • Architectural Shift: The system moves from a descriptive prompting paradigm to a data-constrained paradigm, ensuring the the AI’s output is bounded by verifiable psychological parameters, not just narrative intent.

2. Dynamic Behavioral Consistency Loop (Self-Validation)

  • Improvement: Implementation of continuous self-validation using quantitative metrics (e.g., Root Mean Squared Prediction Error - RMSPE and Inter-rater Reliability - IRR).

  • Specific Action: After generating a response, the system runs a secondary evaluation module that measures the generated text against the expected trajectory defined by the 5-tuple vector. If a response drifts significantly from its assigned personality profile (i.e., high RMSPE), it is flagged for re-prompting or correction before being presented to an operator.

  • Mitigation: This directly addresses the LLM’s tendency to hallucinate or drift into generic responses, ensuring that the character’s behavior remains consistently aligned with its fundamental psychological makeup throughout the a game session.

3. Adaptive Narrative Mapping (Evolving Persona)

  • Improvement: The system is designed to treat psychometric values as dynamic variables, not static constants.

  • Specific Action: As the NPC interacts with the game environment, external events (e g, witnessing a highly stressful event) are fed back into the AC model. The system then calculates how those events should impact the core 5-tuple values (e g, a high Neuroticism score combined with stress causes a temporary spike in N). The LLM is prompted to reflect’s this evolving internal state.

  • Data Flow: This allows for the generation of dynamic character transformations, where the NPC's personality is not just fixed, but genuinely shaped by its experiences within a narrative arc.


The implementation of this framework enables the AI system to achieve the following specific capabilities:

1. Unprecedented Behavioral Fidelity (High Fidelity)

  • Outcome: The AI can reliably maintain a character’s psychological consistency across thousands of interactions, achieving an IRR score close to 1.0. The player will perceive a character whose reactions are predictable based on its defined personality type, but nuanced enough that the responses feel genuine.

2. Quantifiable Depth and Immersion

  • Outcome: Developers gain a precise, quantitative tool for assessing the quality of character design. Instead of subjective feedback (This NPC feels dull), they can objectively measure whether the LLM’s output meets a target profile (e g, The generated dialogue achieved 90% alignment with the 'Highly Agreeable' profile).

3. Context-Aware Narrative Consistency (Dynamic Storytelling)

  • Outcome: The system can generate complex, multi-step narratives where the NPC’s decisions are logically traceable to its internal psychology. For instance, a highly Conscientious NPC will prioritize task completion even over social interaction, and this prioritization will be consistently reflected in its dialogue and actions.

4. Efficient Design Iteration

  • Outcome: By using the 5-tuple vector as a baseline input, designers can rapidly prototype hundreds of distinct personality types without needing to write thousands of lines of descriptive lore, significantly accelerating the development cycle for complex character rosters.

Sources

Related papers