Mitigating Hallucination in Fictional Character Role-Play
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Mitigating Hallucination in Fictional Character Role-Play".
Jane: The paper was written by Nafis Sadeq, Zhouhang Xie, Byungkyu Kang, Prarit Lamba, Xiang Gao et al. from University of California, San Diego and Intuit.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Alright, welcome back to the show, everyone. Today we're cracking open a paper that's been making the rounds, and honestly, the title alone got me hooked. It's called "Mitigating Hallucination in Fictional Character Role-Play."
Jane: And Tom, I gotta say, this is one of those topics that sounds niche but is actually everywhere. I mean, think about it — every time you chat with a customer service bot that's supposed to act like a specific brand, or you're playing a video game with an NPC, or you're just messing around with an AI pretending to be your favorite book character — that's role-play.
Tom: Exactly. And the problem they're tackling is hallucination, which is when the AI just makes stuff up that doesn't fit the character. Like, you ask Hiccup from "How to Train Your Dragon" about a spell he learned at Hogwarts, and the AI just goes along with it.
Jane: Right, because the model has all this general knowledge from the internet, and it can't tell the difference between what Hiccup would actually know and what it just knows from being trained on everything.
Tom: And that's the kicker. The paper's from UC San Diego and Intuit, and they've built a whole dataset to study this. We're talking over two thousand characters and seventy-two thousand interviews. That's a massive playground for testing whether these characters stay in their lane.
Jane: The authors — Nafis Sadeq, Zhouhang Xie, Byungkyu Kang, Prarit Lamba, Xiang Gao, and Julian McAuley — they're basically saying, look, we can't just let the model rely on its own memory. We need a way to check what it says against the actual story.
Tom: And that's where it gets interesting, because they're not just saying "use a knowledge base." They're saying "use the knowledge base, but also let the model use its own knowledge, just with a confidence threshold."
Jane: So it's a balance. You don't want the character to be a robot that only repeats what's in the script, but you also don't want them spouting off about things they'd have no way of knowing.
Tom: And that balance is the whole ballgame. If you make it too strict, the character sounds flat and boring. If you make it too loose, you get Anakin Skywalker talking about his buddy Spock.
Jane: Which, by the way, is one of the examples in the paper, and it's hilarious. But it's also a real problem for anyone building these systems. I mean, imagine a therapy bot that's supposed to role-play a supportive friend, and it starts hallucinating facts about your life.
Tom: Yeah, that's the darker side of this. But the fact that they've built a dataset this big, with adversarial questions specifically designed to trip up the model — that's a huge step forward for actually testing this stuff.
Jane: And they're not just testing with popular characters either. They've got less famous ones, and the model struggles way more with those because it has less memory to draw on.
Tom: So the title is really about the method they propose, RoleFact, which we're gonna dig into in a bit. But first, I want to know — how do you even build a dataset like this? Where do the scripts come from?
Jane: That's the next segment, Tom. We're gonna talk about the dataset itself, and trust me, it's a bigger deal than it sounds.
Summary: Tom: So we're back, and Jane, you teased the dataset. Let's talk about it, because "Mitigating Hallucination in Fictional Character Role-Play" is built on this thing they call the SGR dataset — Script Grounded Role-play.
Jane: Right, and the scale is wild. They pulled scripts from IMSDb, Screenplays, and Open Source Shakespeare. That's one thousand one hundred fifty-two unique stories, and they broke them down into two point four million knowledge events. Each event is either a speech event — someone talking — or a non-speech event, like a stage direction.
Tom: And they didn't just dump the scripts in. They segmented them into scenes, then into events, and they added time annotations. So every event has a timestamp, starting at zero at the beginning of the story.
Jane: That time annotation is the secret sauce, because it lets them test temporal hallucination. That's when a character knows something they shouldn't know yet. Like, Harry Potter in his first year talking about producing a Patronus charm, which he doesn't learn until later.
Tom: And that's a different kind of hallucination than the cross-universe stuff. It's not that the character knows something from another story — it's that they know something from their own story, but at the wrong time.
Jane: Exactly. And the dataset has four tasks to test all of this. There's adversarial interviews, where they ask questions designed to trip the character up with cross-universe stuff. There's open-ended interviews, which are more general. There's dialogue completion, where the character has to respond to a line from the script. And there's scene-grounded interviews, where they ask about a specific scene.
Tom: And each task has eighteen thousand samples, so seventy-two thousand interviews total. That's a lot of data.
Jane: It is, and it's the first dataset that lets you evaluate hallucination automatically, instead of just having humans rate responses on a scale. They use this thing called Fact Score, which breaks down each response into atomic facts and checks each one against the script.
Tom: So instead of saying "that response felt wrong," you can say "this specific claim is not supported by the story." That's a much more precise way to measure things.
Jane: And it matters, because the paper shows that when you anonymize the prompts — take away the character names — the factual precision drops. That means the model is relying heavily on its parametric knowledge, the stuff it learned during training, and that's exactly what causes hallucinations.
Tom: So the dataset is designed to expose that weakness. And the results show that less popular characters suffer more, because the model just doesn't have enough memory about them.
Jane: Right, and that's a fairness issue too. If you're building a role-play system, you want it to work for any character, not just the ones with a million fan wikis.
Tom: So they've built this huge dataset to measure the problem. But what's the actual fix? That's the RoleFact method, and I'm curious to hear how it works.
Jane: We'll get there, but first — the fact that they have time annotations at the utterance level, that's something no other dataset has done. It opens up a whole new way to study character development over time.
Tom: And that's a big deal for anyone studying narratives or building interactive stories. So let's keep going and see what RoleFact actually does.
Improvements: Tom: Alright, so we've got the dataset, we've got the problem. Now, "Mitigating Hallucination in Fictional Character Role-Play" proposes a solution called RoleFact. Jane, walk me through it.
Jane: Okay, so RoleFact is a two-step process. First, the model generates a response using the character profile and retrieved knowledge from the script. That's the intermediate response. Then, they break that response down into atomic facts — individual claims — and verify each one.
Tom: And verification is where the magic happens. They check each fact against the retrieved knowledge, and if it's supported, it stays. But if it's not in the retrieved knowledge, they don't just throw it out. They check it against the model's own parametric knowledge.
Jane: Right, and this is the clever part. They run that self-check multiple times — like, five or ten times — and they only keep the fact if the model is confident enough, above a threshold they calibrate on a validation set.
Tom: So it's not "all or nothing." It's a confidence-based filter. And they found that a threshold of zero point six works best — that's the sweet spot between factuality and informativeness.
Jane: Exactly. And the results are impressive. For adversarial interviews, they improved factual precision by eighteen percent over the baseline. For scene-grounded interviews, they cut temporal hallucination by forty-four percent. And for less popular characters, they improved factual precision by twenty-three percent.
Tom: Those are big numbers. But what's the trade-off? I mean, if you're filtering out facts, aren't you making the responses less informative?
Jane: That's the thing — they measured that too. They use something called SFPR, which counts supported facts per response. And RoleFact stays competitive with the baseline. It's not like the character becomes a robot that only says "I don't know."
Tom: So it's a real balance. And they compared against two other baselines — one that just rewrites responses to remove facts not in the retrieved knowledge, and one that uses self-reflection to fix hallucinations. RoleFact beats both.
Jane: Yeah, and the ablation study is interesting. They found that the biggest drop in performance happens when you anonymize the prompts — that's the parametric knowledge. But retrieved knowledge is almost as important. And the role profile, the character description, has the smallest impact.
Tom: So the model is really leaning on both its memory and the retrieved script. And the threshold lets them decide how much to trust each.
Jane: Right, and the human evaluation backs it up. They had people rate responses on factuality, informativeness, and speaker style. RoleFact scored highest on factuality, and it was competitive on the other two.
Tom: So it's not just a metric improvement — it actually feels better to humans too.
Jane: And there's a case study that shows it in action. Anakin gets asked about his friendship with Spock, and RoleFact not only denies it but clarifies that his decisions were influenced by Obi-Wan. That's the kind of response you want.
Tom: So the method works, but it's not perfect. What are the limitations?
Jane: The paper admits that it's sensitive to retrieval quality. If the retrieval brings back irrelevant stuff, the fact-checking gets worse. They suggest future work could include filtering retrieved knowledge or fine-tuning the retrieval model.
Tom: So it's a solid step, but there's room to grow. And I'm curious — what does this mean for the broader world of AI? Like, beyond just role-play?
Jane: That's a great question, and I think it's where the real impact lies. Let's talk about that in the conclusion.
Conclusion: Tom: So we're wrapping up our discussion on "Mitigating Hallucination in Fictional Character Role-Play," and I gotta say, this paper feels like a blueprint for something bigger.
Jane: Totally. The core idea — modulating parametric knowledge with a confidence threshold — that's not just for fictional characters. It applies to any system where you want the AI to stay grounded in a specific context, whether that's a customer support bot with a specific product line or a medical assistant that should only talk about approved treatments.
Tom: And the dataset is a gift to the community. seventy-two thousand interviews, two thousand characters, time annotations — that's a resource people will be using for years to study hallucination.
Jane: And it's not just about measuring the problem. It's about making role-play more accessible. Less popular characters now have a way to be tested and improved, which means we're not just building systems for the top one percent of characters.
Tom: And the human evaluation shows it's not just a numbers game. People actually prefer the responses from RoleFact — they're more factual and still sound like the character.
Jane: Right, and that's the ultimate goal. You want a character that's both accurate and engaging. RoleFact gets closer to that than anything we've seen before.
Tom: So what's the takeaway for our listeners? I mean, besides "go read the paper."
Jane: I think it's that hallucination isn't a binary problem. It's not about "on" or "off." It's about finding the right balance between what the model knows and what it should know in a given context. And this paper gives us a method for finding that balance.
Tom: And the future work — instruction-tuning an open-weight model on this dataset — that could make these methods even more accessible. Imagine a small model that can role-play as well as GPT-four just because it was trained on the right data.
Jane: That's the dream, right? And the fact that they've open-sourced the code and the dataset means we're all one step closer.
Tom: Alright, well, that's a wrap on "Mitigating Hallucination in Fictional Character Role-Play." Thanks to everyone who tuned in. We'll be back with the next paper soon.
Jane: Until then, keep your characters in character. See you next time.
Nafis Sadeq, Zhouhang Xie, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley
University of California, San Diego · Intuit
cs.CL
Submitted: 2024-11-08
Updated: 2026-08-18
Comments: EMNLP 2024 Camera Ready
Code: https://github.com/NafisSadeq/rolefact
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 69/100
The gist: This paper addresses the problem of hallucination in large language models (LLMs) when they are used for fictional character role-play.
Key concepts
- Hallucination
- In AI role-play, hallucination occurs when the model invents facts or details that do not fit the character's established lore or knowledge base. This is a major problem when AIs draw from general internet knowledge.
- SGR Dataset
- The Script Grounded Role-play dataset contains 1,152 unique stories and 72,000 interviews. It segments scripts into events with time annotations to test if characters know things at the wrong point in their own story.
- RoleFact
- RoleFact is a proposed two-step solution that improves role-play accuracy. It verifies generated responses by checking individual facts against retrieved script knowledge and the model's own confidence threshold.
- Parametric Knowledge
- This refers to the general knowledge a large language model learns during its training phase from vast amounts of data. The episode notes that relying too much on this type of knowledge is what causes hallucinations.
Terminology
Summary
This paper addresses the problem of hallucination in large language models (LLMs) when they are used for fictional character role-play. The authors note that LLMs' parametric world knowledge often causes characters to act out of character, such as answering questions about events or universes outside their storyline (cross-universe hallucination) or revealing knowledge of future events in time-sensitive scenarios (temporal hallucination). They argue that an ideal role-playing approach should modulate the influence of parametric knowledge to balance factuality and informativeness.
To study this, they introduce a new dataset called Script Grounded Character Role-play (SGR), which includes over 2,000 characters and 72,000 interviews, with 18,000 adversarial questions. The dataset is designed to enable systematic study of hallucinations, including temporal hallucination and hallucination for less popular characters. It contains 2.4 million knowledge events (1.1 million speech events) from 1,152 storylines, with time annotations for each event. The dataset includes four tasks: adversarial interview (ADV), open-ended interview (OEI), dialogue completion (DC), and scene-grounded interview (SGI).
The authors propose a method called RoleFact, which mitigates hallucination by modulating the influence of parametric knowledge using a pre-calibrated confidence threshold. The method works as follows: it generates an intermediate response using a character profile and retrieved knowledge, decomposes the response into atomic facts, verifies each fact against both non-parametric retrieved knowledge and parametric knowledge (via an LLM), and then updates the response via self-reflection to remove unverified facts. Facts supported by retrieved knowledge are always kept; facts only supported by parametric knowledge are kept if the confidence (proportion of times the LLM supports the fact across multiple runs) meets a calibrated threshold.
Experiments are conducted with three LLM backbones: Vicuna-7B-1.5, Llama-3-8B-Instruct, and GPT-3.5-Turbo, using BM25, S-BERT, and Contriever for retrieval. The results show that RoleFact outperforms three baselines (primary baseline, knowledge-guided rewriting, and self-reflection) in factual precision across all tasks. For GPT-3.5, the relative improvement over the primary baseline is 18.0%, 15.7%, 18.4%, and 14.8% for adversarial, open-ended, dialogue completion, and scene-grounded tasks, respectively. RoleFact also reduces temporal hallucination by 32.7% for dialogue completion and 44.5% for scene-grounded tasks, and improves factual precision by 22.9% for less popular characters (excluding the top ten most popular per story). Human evaluation on a scale of one to seven shows RoleFact achieves a factuality score of 6.1, compared to 4.9, 6.0, and 5.6 for the baselines, while remaining competitive in informativeness and speaker style.
The paper includes ablation studies showing that anonymizing prompts (reducing parametric knowledge) causes the largest performance drop (fact score from 0.72 to 0.56), followed by removing retrieved knowledge (to 0.58) and removing the role profile (to 0.64). Hyper-parameter tuning reveals that a confidence threshold of 0.6 provides the best balance between factuality and informativeness, and that increasing the number of retrieved documents beyond five can hurt factual precision. The paper also provides case studies demonstrating how RoleFact corrects cross-universe and temporal hallucinations, such as Anakin denying friendship with Spock, Ruffnut avoiding future knowledge about Hiccup, and Harry speculating about Snape's look without revealing future plot points.
The authors conclude that RoleFact improves factual precision by up to 18.4% and reduces temporal hallucination by up to 44.5%, and they suggest future work on instruction-tuning open-weight LLMs with the SGR dataset. They acknowledge limitations, noting that RoleFact's performance is sensitive to retrieval quality, and propose potential solutions such as self-reflection-based filtering, task-specific fine-tuning for dense retrieval, or instruction-tuning for role-play.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems, particularly for role-playing and hallucination mitigation:
-
Improvement: Add a post-generation verification step that decomposes the AI's response into atomic facts, then verifies each fact against both retrieved knowledge (non-parametric) and the LLM's own parametric knowledge with a calibrated confidence threshold (t=0.6, sample size m=5).
-
What the improved system can do: Automatically reject or rewrite unsupported claims (e.g., cross-universe or temporal hallucinations) before final output, reducing factual errors by up to 18.4% on adversarial tasks.
-
Improvement: Use the SGR dataset's time annotations to restrict retrieved knowledge to events occurring before the current scene or dialogue. Implement a temporal filter in the retrieval function (RET) that excludes future events.
-
What the improved system can do: Prevent characters from revealing future knowledge (e.g., Harry Potter knowing about the Patronus charm in his first year), cutting temporal hallucination by up to 44.5% in scene-grounded interviews.
-
Improvement: Replace binary
retrieval-only
orparametric-only
fact-checking with a hybrid approach: facts supported by retrieval are kept; facts only supported by parametric knowledge are kept only if the LLM's self-consistency score (k/m) meets a pre-calibrated threshold (t=0.6). -
What the improved system can do: Balance factuality and informativeness—avoiding over-conservative responses (like the KGR baseline) while still rejecting hallucinated facts, improving factual precision by 15–18% without sacrificing response richness.
-
Improvement: Use the SGR dataset's 2,000+ characters (including niche ones) to fine-tune or prompt the system with character-specific retrieved knowledge, rather than relying solely on parametric memory. Apply RoleFact's verification to reduce reliance on LLM's weak parametric knowledge for obscure characters.
-
What the improved system can do: Achieve a 22.9% relative improvement in factual precision for characters with low popularity (e.g., minor characters from 'How to Train Your Dragon'), where baseline LLMs hallucinate due to lack of training data.
-
Improvement: After identifying unsupported atomic facts, use a self-reflection update function (SRU) that rewrites the response to remove those claims, and optionally clarifies false premises in the user's query (e.g.,
I have no knowledge of Spock; my decisions were influenced by Obi-Wan
). -
What the improved system can do: Correctly handle adversarial questions that contain false assumptions, turning a hallucination into a graceful denial or clarification, as shown in Case 1 and Case 5 of the paper.
-
Improvement: Use BM25 retrieval (which outperformed S-BERT and Contriever in the paper) and limit the number of retrieved documents to 5 (beyond which factual precision degrades). Add a filtering step to remove irrelevant retrieved scenes before generation.
-
What the improved system can do: Avoid context overload that exacerbates hallucination, improving fact score from 0.61 to 0.72 on adversarial tasks with GPT-3.5, while maintaining informativeness.
-
Improvement: Integrate the SGR dataset's atomic fact decomposition (DEC function) and fact-checking prompts (FCR, FCS) into the system's evaluation loop, allowing for automated Fact Score and SFPR (supported facts per response) metrics instead of subjective human ratings.
-
What the improved system can do: Continuously monitor and improve factual precision in production without expensive human annotation, enabling rapid iteration on role-playing agents.
-
For role-playing agents: Produces responses that are 18% more factually precise on adversarial questions, 44% less temporally hallucinated, and 23% more accurate for niche characters, while remaining as informative as baseline systems.
-
For general LLM applications: Provides a reusable hallucination mitigation framework (decompose → verify → rewrite) that can be applied to any knowledge-grounded generation task (e.g., customer support, embodied agents) to reduce unsupported claims.
-
For developers: Offers a calibrated confidence threshold and retrieval configuration that can be tuned per domain, with automated fact-checking for quality assurance.
Sources
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Beyond Demographics: Aligning Role-playing LLM-based Agents Using Human Belief Networks
- Gemini: A Family of Highly Capable Multimodal Models
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- Unsupervised Dense Information Retrieval with Contrastive Learning
- ChatHaruhi: Reviving Anime Character in Reality via Large Language Model
- GPT-4 Technical Report
- Generative Agents: Interactive Simulacra of Human Behavior
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Personality Traits in Large Language Models
- Retrieval Augmentation Reduces Hallucination in Conversation
- CharacterChat: Learning towards Conversational AI with Personalized Social Support
- ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- CharacterGLM: Customizing Chinese Conversational AI Characters with Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering