CHARM: Character Hallucination for Multicultural Role Play Benchmark
cs.CL, cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: 16 pages, 1 figure. Accepted to Findings of EMNLP 2026
Code: https://github.com/sunkyoung19/CHARM
License: http://creativecommons.org/licenses/by/4.0/
The gist: Role-playing large language models (LLMs) are expected to adopt a character's style while also respecting that character's knowledge boundaries.
Terminology
Abstract
Role-playing large language models (LLMs) are expected to adopt a character's style while also respecting that character's knowledge boundaries. Prior evaluations detect character hallucination but rarely distinguish whether errors arise from failure to recognize a boundary or from failure to comply despite recognition. We introduce CHARM, a multicultural benchmark of 40 real and fictional characters drawn from five cultural-linguistic regions, and validated by native reviewers. It probes two boundary types, Temporal (historical vs. modern) and Cross-Universe (entities outside a character's narrative or historical universe), using abstention-enabled multiple-choice questions. We propose a two-stage evaluation that separates Boundary-Awareness (explicit recognition that a query is out of scope) from Boundary-Compliance (abstention when answering concrete questions). Evaluations across six LLMs show that hallucination is driven predominantly by compliance failures. Models frequently acknowledge that a query lies outside the character's knowledge yet still provide factual, out-of-character answers. By re-posing the same questions to the target character, we confirm that a large fraction of these cases are verified parametric overrides; the model stores the relevant fact but fails to suppress it. We also observe systematic cultural variation in these failures, consistent with imbalances in how characters from different regions are represented in model knowledge.
Sources
- Unsupervised Enrichment of Persona-grounded Dialog with Background Stories
- Concept Incongruence: An Exploration of Time and Death in Role Playing
- Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration
- Gemma 3 Technical Report
- GPT-4 Technical Report
- Qwen3 Technical Report
- OpenAI GPT-5 System Card
- RoleBreak: Character Hallucination as a Jailbreak Attack in Role-Playing Systems
- RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering