From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "From Personas to Plot".
Jane: A unified framework for long-form narrative generation and verification was introduced to address the challenges of narrative consistency and plot discontinuity in large language models.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, to wrap up where we are, the core idea of "From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives" is that existing AI fiction generation struggles with maintaining a solid plot line over extended periods. The authors claim their unified framework, which includes Magnet and Atlas, addresses this by using persona-grounded character agents proposing actions based on a shared world state and evolving story goals.
Jane: Exactly, Tom; the main point they're making is that you can generate structurally coherent long-form narratives if you explicitly track the world state and guide the generation through these goal-driven interactions. They show that this approach produces better results when compared to using a single model for prompting or other existing methods like IBSEN.
Lu: It’s interesting how they set up the system where each character agent uses a specific persona—listing things like relationships, goals, and abilities—to decide what action to take next based on the current story goal. That level of detail in defining those agents is what makes their system so robust for storytelling.
Meng: I'm focusing on that shared world state aspect; how they handle the updates between character actions seems critical for preventing narrative drift as the story gets longer. If that state tracking isn't reliable, the whole thing falls apart quickly during generation.
Lalam: From my perspective as an AI, what really stands out is the combination of Magnet for generation and Atlas for verification; having a dedicated pipeline to check scene-level world representations against previous scenes addresses the consistency issue head-on.
Tom: And that brings us right into why it matters: they’re showing how explicit world-state tracking and goal-driven multi-agent generation can actually lead to narratives that are structurally coherent, which is a big step for long stories.
Conclusion: Tom: So, looking at the full picture of "From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives," we have authors like Aayush Aluru, Chloe Ho, Muhammad Hammouri, Myra Malik, Ryan Lagasse, and Arjun Bahuguna on this work. The implication here is that we might finally move closer to creating AI systems capable of producing genuinely sustained and complex fictional narratives rather than just short snippets.
Jane: I think the paper suggests that the future of long-form AI content generation lies in these multi-agent setups, where consistency isn't something the model just guesses, but something that is actively managed through defined character behaviors and updated world rules. It’s about controllability in story creation.
Lu: I see this opening for really wild things; imagine building an entire ecosystem of interconnected narratives where different agent personas constantly negotiate and evolve their goals together across massive story spans. The creative potential for emergent plot structures is immense because the constraints are built into the agents themselves.
Meng: I wonder how we translate this controlled generation into real-world applications; if we can reliably generate coherent long documents, that opens up possibilities for complex procedural content or even detailed simulation scripting where narrative flow matters a lot.
Lalam: I think the most profound impact could be in how AI tools are used to build training data for creative industries; having a system that can reliably produce high-quality, consistent long-form content would really help refine what good storytelling looks like in an AI context.
Tom: It seems the big picture here is that we're not just improving the length of text; we’re improving its internal logic and structural integrity by making the generation process goal-driven and character-aware, which is a solid foundation for future work.
Princeton University · University of Michigan
cs.CL, cs.AI, cs.MA
Submitted: 2026-07-01
Updated: 2026-10-03
Code: https://github.com/OpenDFM/ibsen
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 78/100
The gist: A unified framework for long-form narrative generation and verification was introduced to address the challenges of narrative consistency and plot discontinuity in large language models.
Key concepts
- Character Agents
- These are individual AI personas representing characters in the story, each with a defined personality, goals, and abilities. They use a specialized model to suggest actions relevant to the current story goal and world state.
- Shared World State
- The story's current situation is tracked as a directed graph linking character nodes and state variables. This state is continuously updated using an overwrite conflict resolution strategy to maintain narrative consistency across all agents.
- Atlas Pipeline
- This is a graph-based evaluation system designed to find plot holes or hallucinations. It compares the world state representations of different scenes to detect inconsistencies between what happened previously and what is being described in the current scene.
Terminology
Summary
A unified framework for long-form narrative generation and verification was introduced to address the challenges of narrative consistency and plot discontinuity in large language models. The core contribution is Magnet, a multi-agent goaldriven narrative engine that generates stories with persona-grounded character agents proposing actions based on a shared world state and evolving story goals, complemented by Atlas, a graph-based pipeline designed to detect hallucinations by comparing scene-level world representations. These systems demonstrate that explicit world-state tracking and goal-driven multi-agent generation can produce structurally coherent long-form narratives.
Magnet Framework Components
The Magnet framework is a multi-agent action-critic-narrator generation system designed for long-range coherence. It incorporates several key mechanisms to maintain narrative structure:
-
Character Agents: Each character is defined with a persona containing attributes such as
relationships, personality, goals, character description, role, location, and abilities.
These agents use a DPO-tuned Gemma-4-31B-it model to generate an action based on the current story goal and world state. -
Critic Revision: After an action is generated by a character agent, a
Gemini 2.5 Flash DeepMind LLM critic
evaluates the proposed action to determine if it isrelevant, specific, and consistent with both the character and the current scene.
If not, it provides feedback for revision. -
Narrator Prose Writing: After all characters generate an action, an
Opus 4.7 Anthropic narrator
produces the next paragraph of the story by selectively choosing actions best suited for the scene and creatingcoherent story prose using the selected actions.
-
Shared World State: The story state is represented as a directed graph containing
character nodes, state variable nodes, and edges that represent the relationships between the nodes,
which are updated using anoverwrite conflict resolution strategy.
Goal Sequencing and Dynamics
The framework employs sophisticated goal sequencing to prevent narrative stagnation. The process includes:
-
Goal Generation: A high-level goal is established at the start, and an
Opus 4.7 Anthropic goal generator
produces a follow-up goal upon completion. -
Stalled Narrative Handling: To avoid stalled narratives, goals that have not been completed in
15 time steps
are replaced. -
Narrative Diversity: To maintain diversity, the system generates a new goal that
changes the direction of the story into a new domain every 40 steps.
Atlas Hallucination Evaluation Pipeline
Atlas is a graph-based pipeline introduced to detect inconsistencies by comparing world state representations across scenes. The pipeline involves three sequential passes over the story:
-
Scene Decomposition: Decomposing the script into
scene-level event units, representing each as a node with a name, description, and textual evidence drawn from the screenplay.
-
Entity Extraction: Extracting entities in relation to the events they appear in.
-
Relation Extraction: Extracting
the relations connecting those events and entities.
The system then starts at the second scene and proposes hallucinations based on the inconsistencies between the current scene’s world state and text, and previous scenes’ world states,
which are then verified using story text to provide interpretable graph-grounded results.
Empirical Evaluation Results
Empirical evaluations compared Magnet against a standard prompting approach and IBSEN. At 100 pages, Magnet reduced annotations by 41 compared to the single model baseline and 34 compared to IBSEN. Furthermore, Atlas showed significant improvement in hallucination detection: on the 100-page story, Magnet recorded 6 hallucinations, a 50% and 45% decrease
compared to the single model baseline and IBSEN, respectively. Pairwise rubric evaluation further supported these findings, showing that Magnet achieved higher scores than both baselines at each story length and hierarchical level.
Ablation Analysis
A small ablation study using hierarchical editorial evaluation on a 20-page story confirmed the contribution of individual components. Removing elements such as world state updates,
DPO,
dynamic goal updates,
and the critic module
resulted in higher total LLM editorial annotations compared to Magnet, indicating that these mechanisms are crucial for improving narrative coherence. The results demonstrated that Magnet's improvements arise from the critic-guided refinement, DPO adapter, updating world state, and dynamic goal sequencing.
Conclusion
The work introduces Magnet for controllable long-form narrative generation and Atlas for interpretable hallucination evaluation. The findings suggest that long-form narratives can emerge from explicit world-state tracking and goal-driven multi-agent generation,
providing a foundation for future research in this area. The framework is designed to produce controllable and structurally coherent long-form narrative generation.
Limitations
The work notes that the system remains "computationally expensive due to repeated interaction between multiple generation modules.
Improvements for AI systems
Here are specific, actionable improvements to existing AI systems based on the Magnet and Atlas framework described in the paper:
The core improvement lies in shifting from single-pass, text-based generation models to a structured, multi-agent system with explicit world state tracking and graph-based verification.
Here are the specific improvements you can implement:
-
Enhance Narrative Coherence via Goal Sequencing and Dynamic Goal Generation:
-
Implement Character Grounding through DPO-Tuned Agents:
-
Integrate Real-time World State Tracking for Long-Range Consistency:
-
Establish an Interpretable Hallucination Detection Pipeline (Atlas):
These improvements will enable the improved AI system to perform the following specific tasks:
-
Perform high-fidelity, structurally coherent long-form narrative generation (up to 100 pages) where plot lines remain consistent across extended timelines.
-
Generate character actions and dialogue that are deeply consistent with predefined persona attributes (personality, relationships, abilities) by utilizing DPO-tuned models for action selection and critic revision.
-
Maintain complex factual relationships, character states (e.g., injuries, possessions), and sequential goals across the entire story using a structured world state representation (a directed graph).
-
Detect and pinpoint narrative inconsistencies (hallucinations) by comparing the current scene's world state representation against historical scene representations within a formal graph structure, providing interpretable signals on where model failures occur.
Abstract
Although large language models (LLMs) have demonstrated impressive creative fiction generation, they struggle to maintain narrative consistency and coherent plot lines in long-form stories. In this work, we introduce MAGNET, a multi-agent goal-driven narrative engine for storytelling, which generates stories with persona-grounded character agents that propose actions based on a shared structured state and evolving story goals. We evaluate MAGNET on 20 and 100 page stories using LLM based editor annotations and rubric scoring. At both narrative lengths, MAGNET significantly reduces editor annotations and improves rubric scores compared to single-model prompting, StoryBox, and StoryWriter (p<0.05). Ablation experiments also show that the critic module, structured state, and goal sequencing each contribute significantly to performance (p<0.05). These results suggest that long-form narratives can emerge from explicit structured states, critic-based action refinement, and goal-driven multi-agent generation with loosely specified character personas, providing a foundation for controllable and structurally coherent long-form narrative generation.
Sources
- HalluLens: LLM Hallucination Benchmark
- From World-Gen to Quest-Line: A Dependency-Driven Prompt Pipeline for Coherent RPG Generation
- StoryER: Automatic Story Evaluation via Ranking, Rating and Reasoning
- LLM-State: Open World State Representation for Long-horizon Task Planning with Large Language Model
- Of Human Criteria and Automatic Metrics: A Benchmark of the Evaluation of Story Generation
- WoW: Towards a World omniscient World model Through Embodied Interaction
- Understanding World or Predicting Future? A Comprehensive Survey of World Models
- AgentScope: A Flexible yet Robust Multi-Agent Platform
- A Survey on LLM-as-a-Judge
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics
- NarraBench: A Comprehensive Framework for Narrative Benchmarking
- Beyond State Consistency: Behavior Consistency in Text-Based World Models
- Agents' Room: Narrative Generation through Multi-step Collaboration
- FactKG: Fact Verification via Reasoning on Knowledge Graphs
- Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- From Word to World: Can Large Language Models be Implicit Text-based World Models?
- Narrative Theory-Driven LLM Methods for Automatic Story Generation and Understanding: A Survey
- Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering