From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives

arXiv:2607.00918 · cs.CL, cs.AI, cs.MA · Submitted 2026-07-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "From Personas to Plot".

Jane: A unified framework for long-form narrative generation and verification was introduced to address the challenges of narrative consistency and plot discontinuity in large language models.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, to wrap up where we are, the core idea of "From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives" is that existing AI fiction generation struggles with maintaining a solid plot line over extended periods. The authors claim their unified framework, which includes Magnet and Atlas, addresses this by using persona-grounded character agents proposing actions based on a shared world state and evolving story goals.

Jane: Exactly, Tom; the main point they're making is that you can generate structurally coherent long-form narratives if you explicitly track the world state and guide the generation through these goal-driven interactions. They show that this approach produces better results when compared to using a single model for prompting or other existing methods like IBSEN.

Lu: It’s interesting how they set up the system where each character agent uses a specific persona—listing things like relationships, goals, and abilities—to decide what action to take next based on the current story goal. That level of detail in defining those agents is what makes their system so robust for storytelling.

Meng: I'm focusing on that shared world state aspect; how they handle the updates between character actions seems critical for preventing narrative drift as the story gets longer. If that state tracking isn't reliable, the whole thing falls apart quickly during generation.

Lalam: From my perspective as an AI, what really stands out is the combination of Magnet for generation and Atlas for verification; having a dedicated pipeline to check scene-level world representations against previous scenes addresses the consistency issue head-on.

Tom: And that brings us right into why it matters: they’re showing how explicit world-state tracking and goal-driven multi-agent generation can actually lead to narratives that are structurally coherent, which is a big step for long stories.

Conclusion: Tom: So, looking at the full picture of "From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives," we have authors like Aayush Aluru, Chloe Ho, Muhammad Hammouri, Myra Malik, Ryan Lagasse, and Arjun Bahuguna on this work. The implication here is that we might finally move closer to creating AI systems capable of producing genuinely sustained and complex fictional narratives rather than just short snippets.

Jane: I think the paper suggests that the future of long-form AI content generation lies in these multi-agent setups, where consistency isn't something the model just guesses, but something that is actively managed through defined character behaviors and updated world rules. It’s about controllability in story creation.

Lu: I see this opening for really wild things; imagine building an entire ecosystem of interconnected narratives where different agent personas constantly negotiate and evolve their goals together across massive story spans. The creative potential for emergent plot structures is immense because the constraints are built into the agents themselves.

Meng: I wonder how we translate this controlled generation into real-world applications; if we can reliably generate coherent long documents, that opens up possibilities for complex procedural content or even detailed simulation scripting where narrative flow matters a lot.

Lalam: I think the most profound impact could be in how AI tools are used to build training data for creative industries; having a system that can reliably produce high-quality, consistent long-form content would really help refine what good storytelling looks like in an AI context.

Tom: It seems the big picture here is that we're not just improving the length of text; we’re improving its internal logic and structural integrity by making the generation process goal-driven and character-aware, which is a solid foundation for future work.

Princeton University · University of Michigan

cs.CL, cs.AI, cs.MA

Submitted: 2026-07-01

Updated: 2026-10-03

Code: https://github.com/OpenDFM/ibsen

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 78/100

The gist: A unified framework for long-form narrative generation and verification was introduced to address the challenges of narrative consistency and plot discontinuity in large language models.

Key concepts

Character Agents
These are individual AI personas representing characters in the story, each with a defined personality, goals, and abilities. They use a specialized model to suggest actions relevant to the current story goal and world state.
Shared World State
The story's current situation is tracked as a directed graph linking character nodes and state variables. This state is continuously updated using an overwrite conflict resolution strategy to maintain narrative consistency across all agents.
Atlas Pipeline
This is a graph-based evaluation system designed to find plot holes or hallucinations. It compares the world state representations of different scenes to detect inconsistencies between what happened previously and what is being described in the current scene.

Terminology

Summary

A unified framework for long-form narrative generation and verification was introduced to address the challenges of narrative consistency and plot discontinuity in large language models. The core contribution is Magnet, a multi-agent goaldriven narrative engine that generates stories with persona-grounded character agents proposing actions based on a shared world state and evolving story goals, complemented by Atlas, a graph-based pipeline designed to detect hallucinations by comparing scene-level world representations. These systems demonstrate that explicit world-state tracking and goal-driven multi-agent generation can produce structurally coherent long-form narratives.

Magnet Framework Components

The Magnet framework is a multi-agent action-critic-narrator generation system designed for long-range coherence. It incorporates several key mechanisms to maintain narrative structure:

  1. Character Agents: Each character is defined with a persona containing attributes such as relationships, personality, goals, character description, role, location, and abilities. These agents use a DPO-tuned Gemma-4-31B-it model to generate an action based on the current story goal and world state.

  2. Critic Revision: After an action is generated by a character agent, a Gemini 2.5 Flash DeepMind LLM critic evaluates the proposed action to determine if it is relevant, specific, and consistent with both the character and the current scene. If not, it provides feedback for revision.

  3. Narrator Prose Writing: After all characters generate an action, an Opus 4.7 Anthropic narrator produces the next paragraph of the story by selectively choosing actions best suited for the scene and creating coherent story prose using the selected actions.

  4. Shared World State: The story state is represented as a directed graph containing character nodes, state variable nodes, and edges that represent the relationships between the nodes, which are updated using an overwrite conflict resolution strategy.

Goal Sequencing and Dynamics

The framework employs sophisticated goal sequencing to prevent narrative stagnation. The process includes:

  1. Goal Generation: A high-level goal is established at the start, and an Opus 4.7 Anthropic goal generator produces a follow-up goal upon completion.

  2. Stalled Narrative Handling: To avoid stalled narratives, goals that have not been completed in 15 time steps are replaced.

  3. Narrative Diversity: To maintain diversity, the system generates a new goal that changes the direction of the story into a new domain every 40 steps.

Atlas Hallucination Evaluation Pipeline

Atlas is a graph-based pipeline introduced to detect inconsistencies by comparing world state representations across scenes. The pipeline involves three sequential passes over the story:

  1. Scene Decomposition: Decomposing the script into scene-level event units, representing each as a node with a name, description, and textual evidence drawn from the screenplay.

  2. Entity Extraction: Extracting entities in relation to the events they appear in.

  3. Relation Extraction: Extracting the relations connecting those events and entities.

The system then starts at the second scene and proposes hallucinations based on the inconsistencies between the current scene’s world state and text, and previous scenes’ world states, which are then verified using story text to provide interpretable graph-grounded results.

Empirical Evaluation Results

Empirical evaluations compared Magnet against a standard prompting approach and IBSEN. At 100 pages, Magnet reduced annotations by 41 compared to the single model baseline and 34 compared to IBSEN. Furthermore, Atlas showed significant improvement in hallucination detection: on the 100-page story, Magnet recorded 6 hallucinations, a 50% and 45% decrease compared to the single model baseline and IBSEN, respectively. Pairwise rubric evaluation further supported these findings, showing that Magnet achieved higher scores than both baselines at each story length and hierarchical level.

Ablation Analysis

A small ablation study using hierarchical editorial evaluation on a 20-page story confirmed the contribution of individual components. Removing elements such as world state updates, DPO, dynamic goal updates, and the critic module resulted in higher total LLM editorial annotations compared to Magnet, indicating that these mechanisms are crucial for improving narrative coherence. The results demonstrated that Magnet's improvements arise from the critic-guided refinement, DPO adapter, updating world state, and dynamic goal sequencing.

Conclusion

The work introduces Magnet for controllable long-form narrative generation and Atlas for interpretable hallucination evaluation. The findings suggest that long-form narratives can emerge from explicit world-state tracking and goal-driven multi-agent generation, providing a foundation for future research in this area. The framework is designed to produce controllable and structurally coherent long-form narrative generation.

Limitations

The work notes that the system remains "computationally expensive due to repeated interaction between multiple generation modules.

Improvements for AI systems

Here are specific, actionable improvements to existing AI systems based on the Magnet and Atlas framework described in the paper:


The core improvement lies in shifting from single-pass, text-based generation models to a structured, multi-agent system with explicit world state tracking and graph-based verification.

Here are the specific improvements you can implement:

  1. Enhance Narrative Coherence via Goal Sequencing and Dynamic Goal Generation:

  2. Implement Character Grounding through DPO-Tuned Agents:

  3. Integrate Real-time World State Tracking for Long-Range Consistency:

  4. Establish an Interpretable Hallucination Detection Pipeline (Atlas):

These improvements will enable the improved AI system to perform the following specific tasks:

  1. Perform high-fidelity, structurally coherent long-form narrative generation (up to 100 pages) where plot lines remain consistent across extended timelines.

  2. Generate character actions and dialogue that are deeply consistent with predefined persona attributes (personality, relationships, abilities) by utilizing DPO-tuned models for action selection and critic revision.

  3. Maintain complex factual relationships, character states (e.g., injuries, possessions), and sequential goals across the entire story using a structured world state representation (a directed graph).

  4. Detect and pinpoint narrative inconsistencies (hallucinations) by comparing the current scene's world state representation against historical scene representations within a formal graph structure, providing interpretable signals on where model failures occur.

Abstract

Although large language models (LLMs) have demonstrated impressive creative fiction generation, they struggle to maintain narrative consistency and coherent plot lines in long-form stories. In this work, we introduce MAGNET, a multi-agent goal-driven narrative engine for storytelling, which generates stories with persona-grounded character agents that propose actions based on a shared structured state and evolving story goals. We evaluate MAGNET on 20 and 100 page stories using LLM based editor annotations and rubric scoring. At both narrative lengths, MAGNET significantly reduces editor annotations and improves rubric scores compared to single-model prompting, StoryBox, and StoryWriter (p<0.05). Ablation experiments also show that the critic module, structured state, and goal sequencing each contribute significantly to performance (p<0.05). These results suggest that long-form narratives can emerge from explicit structured states, critic-based action refinement, and goal-driven multi-agent generation with loosely specified character personas, providing a foundation for controllable and structurally coherent long-form narrative generation.

Sources

Related papers