PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue
Bo Tang, Jianan Yang, Junyi Zhu, Yiquan Wu, Rui Zhao, Zhengyu Yang, Yang Zhang, Feiyu Xiong, Zhiyu Li, Jiajun Shen
MemTensor (Shanghai) Technology · KU Leuven · Zhejiang University · University of Chinese Academy of Sciences · Sinar Mas Paper (China) Investment Company Limited · The Hong Kong Polytechnic University
cs.CL, cs.AI
Submitted: 2026-08-07
Updated: 2026-08-10
Code: https://github.com/MemTensor/PHASE-Tree
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: The paper, authored by Bo Tang, Jianan Yang, Junyi Zhu, Yiquan Wu, Rui Zhao, Zhengyu Yang, Yang Zhang, Feiyu Xiong, Zhiyu Li, and Jiajun Shen, addresses the problem of stale-state failure in
Terminology
Summary
The paper, authored by Bo Tang, Jianan Yang, Junyi Zhu, Yiquan Wu, Rui Zhao, Zhengyu Yang, Yang Zhang, Feiyu Xiong, Zhiyu Li, and Jiajun Shen, addresses the problem of stale-state failure in long-horizon role-playing dialogue. The authors state: "Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative. Yet existing work falls short on two fronts: representations are typically static profiles that cannot be updated locally without destabilizing unchanged traits, and benchmarks mainly test persona preservation and memory recall rather than whether a model speaks from a character's currently evolved state. We address both."
The paper makes three contributions:
-
PHASE-Tree character-state modeling:
A representation that decomposes character state into immutable identity facts and mutable persona, session, and moment attributes, with cross-episode evolution gated by resistance–evidence–cooldown policies.
-
LongEvoRoleBench: "A benchmark suite that standardizes eight role-playing corpora into a unified next-utterance protocol for evaluating both within-scene and cross-episode character-state evolution, with metrics tied to the current time-t state rather than a frozen profile."
-
Systematic dual-paradigm validation:
We evaluate the same PHASE-Tree state under both explicit textual provision and implicit parametric adaptation, benchmarking against a comprehensive suite of ablation variants and external baselines.
The motivating example is Chandler Bing from Friends: "early on he is sarcastic and commitment-phobic, but by later seasons he has grown into a husband who trusts his partner. A model that still treats commitment as a punchline in a marriage scene sounds superficially like Chandler while speaking from the wrong narrative state—the model has not forgotten his voice, but has forgotten that the character has changed. We call this stale-state failure."
PHASE-Tree (Psychology-grounded Hierarchical Attribute-Structured Evolving Tree) models character state at time t as a four-part structured tree
:
St = ⟨I, St persona, St session, St moment⟩
-
Identity root (I) —
immutable identity facts (name, gender, and backstory)
-
Persona stratum (high resistance) —
long-term dispositions and relatively stable profile attributes (personality, speaking style, behavioral tendencies, hobbies, relationships, occupation, demographics)
-
Session stratum (moderate resistance) —
within-scene characteristic adaptations (newly learned information, attitude shifts, commitments, and stance changes) accumulated during the current scene
-
Moment stratum (low resistance) —
transient state affect, specifically the dominant emotion, its intensity, and the triggering scene context, refreshed at each scene boundary
The persona–session distinction is grounded in McAdams' separation of broad dispositional traits from contextualized characteristic adaptations,
while the moment stratum follows the state–trait distinction in affect psychology.
Every editable field is independently addressable: an update targets one field without rewriting siblings.
An initial raw character profile is mapped once into PHASE-Tree via a fixed zero-shot GPT-4.1 extractor,
applied uniformly across all eight corpora without corpus-specific manual authoring or rule engineering.
Field values remain free-text, and the baseline tree can be cached and reused across inference calls.
Within a scene, an LLM analyzes only the observed prefix ct together with the scene-start identity and persona
to extract a third-person session entry covering newly learned information, attitude shifts, commitments, and stance changes
plus a moment snapshot capturing the dominant emotion, its intensity, and the current scene context.
When the scene closes, the local records are archived as evidence for subsequent persona evolution.
A three-stage pipeline updates persona fields independently, with per-field resistance calibrated to narrative pacing:
-
Stage 1: Evidence Accumulation —
a separate LLM pass scans each scene from the character's perspective, identifies salient session events... and labels each with a significance level (medium or high).
-
Stage 2: Resistance-Gated Judgment — after each episode, an LLM proposes per-field candidate updates, and
a deterministic validator accepts or rejects each proposal under threshold checks.
Three checks must hold jointly:
update(f) ⇐⇒ nep(f) ≥ τ ep r(f) ∧ nhigh(f) ≥ τ high r(f) ∧ Δep(f) ≥ τ cd r(f)
Concretely, "personality and speaking style are core fields (requiring evidence from ≥16 episodes with ≥6 high-significance entries), behavioral tendencies is moderate (≥3 episodes), and relationships, occupation, hobbies, and demographics are low (1 high-significance entry or 2 medium-significance entries suffice...). Thus core traits demand evidence spanning roughly a full season, whereas a relationship status can update from a single decisive event. Thresholds are
manually set once from narrative-pacing priors and held fixed across all four long-dialogue corpora."
- Stage 3: Incremental Field Update — uses either an incremental merge (
adds or refines content while preserving the previous value
) or a replacement merge (substitutes the previous value entirely and is reserved for explicit contradictions
). Post-processing patches handle edge cases:stale relationship entries are demoted when they lack recent evidence, reciprocity gaps between interacting characters are repaired, and continuity is forward-filled to avoid regression across sequential episodes.
Two complementary paradigms consume the state at inference:
-
Explicit Textual Provision:
The tree is serialized into structured natural-language paragraphs (identity facts, persona traits, session adaptations, and momentary affect) and concatenated with the dialogue context
— the primary validated path. -
Implicit Parametric Adaptation:
a hypernetwork Hϕ maps the embedded character state to a LoRA adapter ∆θt = Hϕ(emb(St)) that is merged into the backbone. The prompt then carries only dialogue context
— described asa token-efficient alternative.
The benchmark "comprises eight datasets evaluated under unified metrics: four long-dialogue corpora constitute the core test of cross-episode character evolution, while four short-dialogue corpora serve as a control setting that isolates within-scene consistency and local state tracking without cross-episode evolution."
-
Short-dialogue sources: RAIDEN, CharacterEval (RPCA benchmarks), SimsConv, and ChatHaruhi.
-
Long-dialogue sources:
Friends (ConvoKit), The Office and Star Trek (public episode transcripts), and Harry Potter (HPD), each tracking six main characters whose relationships and affect evolve across seasons or books.
Evaluation uses complementary random and OOD holdouts
: short-dialogue OOD tests select profile outliers via embedding clustering; long-dialogue OOD tests chronologically hold out later seasons, where relationships, beliefs, and affect may have evolved substantially.
Crucially, scoring uses the matching time-t state rather than a frozen profile.
Three metrics are used:
-
Character Score (Char) and Semantic Score (Sem) —
independent 1–5 LLM-as-Judge ratings (GPT-4.1, greedy decoding) for profile consistency and contextual coherence
-
Embedding Score (Emb) —
cosine similarity to the ground-truth response under OpenAI's text-embedding-3-small
The authors omit n-gram metrics (BLEU, ROUGE) because long-horizon role-playing admits many surface-divergent yet equally valid continuations.
The main comparison uses Qwen2.5-7B-Instruct as the shared backbone with temperature 0.3, max 256 tokens, seed 42. Internal ablations are: Base (no profile), RP (raw profile), NR (LLM-rewritten profile), ST (structured tree, frozen), DT (tree with cross-episode persona evolution, no session/moment), and PT (full pipeline). External baselines include RAG, PAG, CFG (textual) and MT-LoRA, Steering, OPPU, P2P (parametric).
Internal ablation (explicit textual provision): PT ranks first on all eight datasets for Sem and Emb and on five of eight for Char, yielding the best score in 21 of 24 dataset–metric cells. On the four long-dialogue corpora, PT leads 11 of 12 cells.
The DT→PT transition is the only structured transition that improves all three metrics (+0.110 Char, +0.334 Sem, and +0.043 Emb), showing that session and moment layers supply transient cues missed by cross-episode evolution alone.
External comparison (textual provision): Ours ranks first in all 12 long-dialogue dataset–metric cells, first on 18 of 24 cells overall, and in the top two on 20.
The improvement over the strongest textual baseline: +0.49 Char (3.00 vs. PAG's 2.51, +19.7%), +0.41 Sem (3.70 vs. RAG's 3.29, +12.4%), and +0.04 Emb (0.31 vs. RAG's 0.27, +15.1%).
Implicit parametric adaptation: Ours ranks first on 8 of 24 dataset–metric cells and in the top two on 18, leading on Sem in both short-dialogue and long-dialogue averages and tying for first on long-dialogue Emb.
However, internal variants are very close on both Sem and Emb (most rows within ±0.01),
and "Char remains lower than under textual provision. This pattern indicates that the bottleneck lies in the profile-to-LoRA mapping, which compresses away the fine-grained state detail that distinguishes tree variants, rather than in the input representation."
Cross-backbone analysis (Qwen3-0.6B, Gemma-4-E4B, Qwen2.5-7B, Qwen3-32B): PT achieves the best long-dialogue Sem on all four backbones.
Token efficiency: "Parametric adaptation eliminates all profile tokens from the prompt (matching the Base cost of 204 short / 372 long tokens), whereas textual provision requires carrying the profile explicitly: 471 tokens on short dialogues... and 1736 on long dialogues... but yields the best long-dialogue Sem and Emb in our comparison."
Human evaluation: In a blinded 200-response study, human ratings correlate with the GPT-4.1 judge (Pearson r = 0.65); on descriptive n = 10 PT and NR prompt subsets, the Overall difference is +0.20.
Across 50,232 matched question IDs, "GPT-4.1 yields pooled Overall ∆ = +0.087 (Wilcoxon signed-rank p < 0.001)."
Judge robustness: The persona-reference ablation shows The two reference conditions produce the same overall Sem and Emb pattern, while Char responds to the lexical content of the profile reference.
The judge-model analysis (GPT-4.1, GLM-5.2, DeepSeek-V4-Flash) finds the Sem conclusion is stable across judges, whereas Char rankings are judge-dependent.
The authors conclude: "We address stale-state failure in long-horizon role-playing. PHASE-Tree separates an immutable identity root from editable persona, session, and moment fields... Under explicit textual provision, our method ranks first on 21 of 24 internal cells and all 12 long-dialogue external-comparison cells... The gap between paradigms points to profile-to-LoRA compression as the bottleneck, not the tree representation itself... Together, these contributions establish evolution-aware role-playing as a first-class subtask alongside persona preservation and memory recall."
Improvements for AI systems
The improved AI system can:
-
Maintain a four-part character state tree — an immutable identity root (name, gender, backstory) plus three mutable strata: persona (personality, speaking style, behavioral tendencies, relationships, occupation, hobbies, demographics), session (within-scene new information, attitude shifts, commitments, stance changes), and moment (dominant emotion, intensity, triggering scene context). Each field is independently addressable, so updating one trait does not rewrite or destabilize unchanged traits.
-
Avoid stale-state failure across long narratives — e.g., a model role-playing Chandler Bing can keep his early sarcastic voice while correctly speaking as a married, commitment-trusting husband in later seasons, because generation is conditioned on the current time- t state rather than a frozen profile.
-
Evolve persona only when evidence justifies it — use a resistance-gated pipeline that accumulates scene-level evidence, labels significance (medium/high), and applies per-field thresholds plus cooldown checks before updating. Core fields like personality and speaking style require evidence from ≥16 episodes with ≥6 high-significance entries; relationships, occupation, hobbies, and demographics can update from one high-significance or two medium-significance events. Updates are either incremental merges (preserving prior content) or replacement merges (reserved for explicit contradictions).
-
Track within-scene transient state — at each scene boundary, extract a third-person session entry (newly learned information, attitude shifts, commitments, stance changes) and a moment snapshot (dominant emotion, intensity, trigger context). These local records condition the next utterance and are later archived as evidence for persona evolution.
-
Support two complementary inference paradigms:
-
Explicit textual provision: serialize the tree into structured natural-language paragraphs and concatenate them with dialogue context for best accuracy.
-
Implicit parametric adaptation: use a hypernetwork to map the embedded character state to a LoRA adapter merged into the backbone, eliminating all profile tokens from the prompt while retaining the evolving state — useful for long-horizon dialogue under tight context budgets.
-
Evaluate with time-t matched state — score responses against the character state at the current narrative time, not a frozen profile, using character consistency, semantic coherence, and embedding similarity metrics, with chronological OOD holdouts that test evolution across seasons.
-
Generalize across backbones and domains — apply the same state representation and update policy to different model sizes (0.6B to 32B) and to eight role-playing corpora, including cross-episode long-dialogue datasets and within-scene short-dialogue controls.
-
Scale to long-horizon interactive applications — games, virtual companions, tutoring agents, and therapeutic chatbots where a user’s relationship, goals, and emotional state change over months of interaction; the system can adapt its behavior to the current stage of the relationship without forgetting core identity or regressing to outdated assumptions.
-
Provide explainable state updates — because each persona field has deterministic acceptance checks and archived evidence, the system can justify why a trait changed (or did not change), enabling user-facing transparency and debuggability in AI storytelling and dialogue agents.
Sources
- Beyond Fixed Psychological Personas: State Beats Trait, but Language Models are State-Blind
- Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes
- ChatHaruhi: Reviving Anime Character in Reality via Large Language Model
- HorizonBench: Long-Horizon Personalization with Evolving Preferences
- SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
- SPASM: Stable Persona-driven Agent Simulation for Multi-turn Dialogue Generation
- PersonaVLM: Long-Term Personalized Multimodal LLMs
- Dynamic Personality Adaptation in Large Language Models via State Machines
- Instant Personalized Large Language Model Adaptation via Hypernetwork
- Steering Language Models With Activation Engineering
- The Need for a Socially-Grounded Persona Framework for User Simulation
- Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
- AdaMem: Adaptive User-Centric Memory for Long-Horizon Dialogue Agents
- Inside Out: Evolving User-Centric Core Memory Trees for Long-Term Personalized Dialogue Systems
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering