Mitigating Hallucination in Fictional Character Role-Play
summary
The gist
This paper addresses the problem of hallucination in large language models (LLMs) when they are used for fictional character role-play.
In short
The episode discusses 'Mitigating Hallucination in Fictional Character Role-Play,' a paper from UC San Diego and Intuit. Hosts review the SGR dataset, which contains 72,000 interviews across 2,4 million events. They examine the RoleFact method for improving factual accuracy by balancing retrieved knowledge with model confidence.
Key concepts
- Hallucination
- In AI role-play, hallucination occurs when the model invents facts or details that do not fit the character's established lore or knowledge base. This is a major problem when AIs draw from general internet knowledge.
- SGR Dataset
- The Script Grounded Role-play dataset contains 1,152 unique stories and 72,000 interviews. It segments scripts into events with time annotations to test if characters know things at the wrong point in their own story.
- RoleFact
- RoleFact is a proposed two-step solution that improves role-play accuracy. It verifies generated responses by checking individual facts against retrieved script knowledge and the model's own confidence threshold.
- Parametric Knowledge
- This refers to the general knowledge a large language model learns during its training phase from vast amounts of data. The episode notes that relying too much on this type of knowledge is what causes hallucinations.
Terminology used across episodes
This episode discusses
- Mitigating Hallucination in Fictional Character Role-Play · Paper Radio
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Beyond Demographics: Aligning Role-playing LLM-based Agents Using Human Belief Networks
- Gemini: A Family of Highly Capable Multimodal Models
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- Unsupervised Dense Information Retrieval with Contrastive Learning
- ChatHaruhi: Reviving Anime Character in Reality via Large Language Model
- GPT-4 Technical Report
- Generative Agents: Interactive Simulacra of Human Behavior
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Personality Traits in Large Language Models
- Retrieval Augmentation Reduces Hallucination in Conversation
- CharacterChat: Learning towards Conversational AI with Personalized Social Support
- ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- CharacterGLM: Customizing Chinese Conversational AI Characters with Large Language Models
The paper
Mitigating Hallucination in Fictional Character Role-Play · Read on arXiv
Nafis Sadeq, Zhouhang Xie, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley
University of California, San Diego · Intuit
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Mitigating Hallucination in Fictional Character Role-Play".
Jane: The paper was written by Nafis Sadeq, Zhouhang Xie, Byungkyu Kang, Prarit Lamba, Xiang Gao et al. from University of California, San Diego and Intuit.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Alright, welcome back to the show, everyone. Today we're cracking open a paper that's been making the rounds, and honestly, the title alone got me hooked. It's called "Mitigating Hallucination in Fictional Character Role-Play."
Jane: And Tom, I gotta say, this is one of those topics that sounds niche but is actually everywhere. I mean, think about it — every time you chat with a customer service bot that's supposed to act like a specific brand, or you're playing a video game with an NPC, or you're just messing around with an AI pretending to be your favorite book character — that's role-play.
Tom: Exactly. And the problem they're tackling is hallucination, which is when the AI just makes stuff up that doesn't fit the character. Like, you ask Hiccup from "How to Train Your Dragon" about a spell he learned at Hogwarts, and the AI just goes along with it.
Jane: Right, because the model has all this general knowledge from the internet, and it can't tell the difference between what Hiccup would actually know and what it just knows from being trained on everything.
Tom: And that's the kicker. The paper's from UC San Diego and Intuit, and they've built a whole dataset to study this. We're talking over two thousand characters and seventy-two thousand interviews. That's a massive playground for testing whether these characters stay in their lane.
Jane: The authors — Nafis Sadeq, Zhouhang Xie, Byungkyu Kang, Prarit Lamba, Xiang Gao, and Julian McAuley — they're basically saying, look, we can't just let the model rely on its own memory. We need a way to check what it says against the actual story.
Tom: And that's where it gets interesting, because they're not just saying "use a knowledge base." They're saying "use the knowledge base, but also let the model use its own knowledge, just with a confidence threshold."
Jane: So it's a balance. You don't want the character to be a robot that only repeats what's in the script, but you also don't want them spouting off about things they'd have no way of knowing.
Tom: And that balance is the whole ballgame. If you make it too strict, the character sounds flat and boring. If you make it too loose, you get Anakin Skywalker talking about his buddy Spock.
Jane: Which, by the way, is one of the examples in the paper, and it's hilarious. But it's also a real problem for anyone building these systems. I mean, imagine a therapy bot that's supposed to role-play a supportive friend, and it starts hallucinating facts about your life.
Tom: Yeah, that's the darker side of this. But the fact that they've built a dataset this big, with adversarial questions specifically designed to trip up the model — that's a huge step forward for actually testing this stuff.
Jane: And they're not just testing with popular characters either. They've got less famous ones, and the model struggles way more with those because it has less memory to draw on.
Tom: So the title is really about the method they propose, RoleFact, which we're gonna dig into in a bit. But first, I want to know — how do you even build a dataset like this? Where do the scripts come from?
Jane: That's the next segment, Tom. We're gonna talk about the dataset itself, and trust me, it's a bigger deal than it sounds.
Summary: Tom: So we're back, and Jane, you teased the dataset. Let's talk about it, because "Mitigating Hallucination in Fictional Character Role-Play" is built on this thing they call the SGR dataset — Script Grounded Role-play.
Jane: Right, and the scale is wild. They pulled scripts from IMSDb, Screenplays, and Open Source Shakespeare. That's one thousand one hundred fifty-two unique stories, and they broke them down into two point four million knowledge events. Each event is either a speech event — someone talking — or a non-speech event, like a stage direction.
Tom: And they didn't just dump the scripts in. They segmented them into scenes, then into events, and they added time annotations. So every event has a timestamp, starting at zero at the beginning of the story.
Jane: That time annotation is the secret sauce, because it lets them test temporal hallucination. That's when a character knows something they shouldn't know yet. Like, Harry Potter in his first year talking about producing a Patronus charm, which he doesn't learn until later.
Tom: And that's a different kind of hallucination than the cross-universe stuff. It's not that the character knows something from another story — it's that they know something from their own story, but at the wrong time.
Jane: Exactly. And the dataset has four tasks to test all of this. There's adversarial interviews, where they ask questions designed to trip the character up with cross-universe stuff. There's open-ended interviews, which are more general. There's dialogue completion, where the character has to respond to a line from the script. And there's scene-grounded interviews, where they ask about a specific scene.
Tom: And each task has eighteen thousand samples, so seventy-two thousand interviews total. That's a lot of data.
Jane: It is, and it's the first dataset that lets you evaluate hallucination automatically, instead of just having humans rate responses on a scale. They use this thing called Fact Score, which breaks down each response into atomic facts and checks each one against the script.
Tom: So instead of saying "that response felt wrong," you can say "this specific claim is not supported by the story." That's a much more precise way to measure things.
Jane: And it matters, because the paper shows that when you anonymize the prompts — take away the character names — the factual precision drops. That means the model is relying heavily on its parametric knowledge, the stuff it learned during training, and that's exactly what causes hallucinations.
Tom: So the dataset is designed to expose that weakness. And the results show that less popular characters suffer more, because the model just doesn't have enough memory about them.
Jane: Right, and that's a fairness issue too. If you're building a role-play system, you want it to work for any character, not just the ones with a million fan wikis.
Tom: So they've built this huge dataset to measure the problem. But what's the actual fix? That's the RoleFact method, and I'm curious to hear how it works.
Jane: We'll get there, but first — the fact that they have time annotations at the utterance level, that's something no other dataset has done. It opens up a whole new way to study character development over time.
Tom: And that's a big deal for anyone studying narratives or building interactive stories. So let's keep going and see what RoleFact actually does.
Improvements: Tom: Alright, so we've got the dataset, we've got the problem. Now, "Mitigating Hallucination in Fictional Character Role-Play" proposes a solution called RoleFact. Jane, walk me through it.
Jane: Okay, so RoleFact is a two-step process. First, the model generates a response using the character profile and retrieved knowledge from the script. That's the intermediate response. Then, they break that response down into atomic facts — individual claims — and verify each one.
Tom: And verification is where the magic happens. They check each fact against the retrieved knowledge, and if it's supported, it stays. But if it's not in the retrieved knowledge, they don't just throw it out. They check it against the model's own parametric knowledge.
Jane: Right, and this is the clever part. They run that self-check multiple times — like, five or ten times — and they only keep the fact if the model is confident enough, above a threshold they calibrate on a validation set.
Tom: So it's not "all or nothing." It's a confidence-based filter. And they found that a threshold of zero point six works best — that's the sweet spot between factuality and informativeness.
Jane: Exactly. And the results are impressive. For adversarial interviews, they improved factual precision by eighteen percent over the baseline. For scene-grounded interviews, they cut temporal hallucination by forty-four percent. And for less popular characters, they improved factual precision by twenty-three percent.
Tom: Those are big numbers. But what's the trade-off? I mean, if you're filtering out facts, aren't you making the responses less informative?
Jane: That's the thing — they measured that too. They use something called SFPR, which counts supported facts per response. And RoleFact stays competitive with the baseline. It's not like the character becomes a robot that only says "I don't know."
Tom: So it's a real balance. And they compared against two other baselines — one that just rewrites responses to remove facts not in the retrieved knowledge, and one that uses self-reflection to fix hallucinations. RoleFact beats both.
Jane: Yeah, and the ablation study is interesting. They found that the biggest drop in performance happens when you anonymize the prompts — that's the parametric knowledge. But retrieved knowledge is almost as important. And the role profile, the character description, has the smallest impact.
Tom: So the model is really leaning on both its memory and the retrieved script. And the threshold lets them decide how much to trust each.
Jane: Right, and the human evaluation backs it up. They had people rate responses on factuality, informativeness, and speaker style. RoleFact scored highest on factuality, and it was competitive on the other two.
Tom: So it's not just a metric improvement — it actually feels better to humans too.
Jane: And there's a case study that shows it in action. Anakin gets asked about his friendship with Spock, and RoleFact not only denies it but clarifies that his decisions were influenced by Obi-Wan. That's the kind of response you want.
Tom: So the method works, but it's not perfect. What are the limitations?
Jane: The paper admits that it's sensitive to retrieval quality. If the retrieval brings back irrelevant stuff, the fact-checking gets worse. They suggest future work could include filtering retrieved knowledge or fine-tuning the retrieval model.
Tom: So it's a solid step, but there's room to grow. And I'm curious — what does this mean for the broader world of AI? Like, beyond just role-play?
Jane: That's a great question, and I think it's where the real impact lies. Let's talk about that in the conclusion.
Conclusion: Tom: So we're wrapping up our discussion on "Mitigating Hallucination in Fictional Character Role-Play," and I gotta say, this paper feels like a blueprint for something bigger.
Jane: Totally. The core idea — modulating parametric knowledge with a confidence threshold — that's not just for fictional characters. It applies to any system where you want the AI to stay grounded in a specific context, whether that's a customer support bot with a specific product line or a medical assistant that should only talk about approved treatments.
Tom: And the dataset is a gift to the community. seventy-two thousand interviews, two thousand characters, time annotations — that's a resource people will be using for years to study hallucination.
Jane: And it's not just about measuring the problem. It's about making role-play more accessible. Less popular characters now have a way to be tested and improved, which means we're not just building systems for the top one percent of characters.
Tom: And the human evaluation shows it's not just a numbers game. People actually prefer the responses from RoleFact — they're more factual and still sound like the character.
Jane: Right, and that's the ultimate goal. You want a character that's both accurate and engaging. RoleFact gets closer to that than anything we've seen before.
Tom: So what's the takeaway for our listeners? I mean, besides "go read the paper."
Jane: I think it's that hallucination isn't a binary problem. It's not about "on" or "off." It's about finding the right balance between what the model knows and what it should know in a given context. And this paper gives us a method for finding that balance.
Tom: And the future work — instruction-tuning an open-weight model on this dataset — that could make these methods even more accessible. Imagine a small model that can role-play as well as GPT-four just because it was trained on the right data.
Jane: That's the dream, right? And the fact that they've open-sourced the code and the dataset means we're all one step closer.
Tom: Alright, well, that's a wrap on "Mitigating Hallucination in Fictional Character Role-Play." Thanks to everyone who tuned in. We'll be back with the next paper soon.
Jane: Until then, keep your characters in character. See you next time.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization