How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling

arXiv:2609.02482 · cs.CL · Submitted 2026-09-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling".

Jane: The paper was written by Katrin Rohrbacher, Björn Nieth, Emmanuelle Salin, Bjoern Eskofier and Michaela Mahlberg from The Text and Language Lab, Department of Digital Humanities and Social Studies, Friedrich-Alexander University (FAU) Erlangen-Nürnberg and Department of Artificial Intelligence in Biomedical Engineering (AIBE), Friedrich-Alexander University (FAU) Erlangen-Nürnberg and Chair of AI-supported Therapy Decisions, Ludwig Maximilian University (LMU) Munich and Department of Linguistics and Communication, University of Birmingham and Munich Center for Machine Learning (MCML) and Institute of AI for Health, Helmholtz Centre Munich.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Jane: So, let's talk about the title again and what it implies for its scope. The paper is "How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling," and it sets up a very specific investigative goal: to measure setting as one of the primary mechanisms of worldbuilding.

Tom: And the implications are massive because, if the AI isn't building worlds like humans do, what we might think is just a stylistic variation could actually be something much deeper about how AI thinks about space and narrative flow.

Lu: It’s not just a surface-level difference; it’s an internal difference in how LLMs prioritize different elements of the environment. The way they structure the storyworld seems fundamentally different from its human counterpart, which is what makes this title so provocative.

Meng: I think the initial comparison between this AI generation and Project Gutenberg texts is very practical for us because it provides a clear historical baseline to measure against, allowing us to see if these models are just reproducing old styles or if they’ are generating something new.

Lalam: This difference suggests that the way we perceive AI-generated content might change over time because its narrative construction is inherently different from our traditional understanding of human fiction.

Tom: It's a strong contrast to be able to measure so clear a distinction, and that sets the stage for seeing exactly where those differences are in the next segment.

Summary: Jane: The paper’s summary shows us the core finding: that LLMs systematically overproduce "perceived space" compared to human authors. This isn't just a tiny difference, which is quite striking given how much we usually rely on a baseline for AI performance.

Tom: We’re seeing that when the human authors use action space—movement and physical interaction—the characters are grounded in their environment, but the LLM-generated texts tend to emphasize atmosphere and feeling instead of concrete action.

Lu: The data is quite specific about this too; for instance, looking at the opening fifteen sentences in Figure two GPT four point one reaches a normalized frequency around zero point four seven in English, which is roughly twice the human baseline—that's a huge overproduction of atmosphere!

Meng: And it’s important to note that this isn't just limited to the opening sentence; the authors show us that this bias persists across narrative time, meaning the AI’s preference for mood is consistent throughout a long story.

Lalam: This means that when we think about cultural impact, LLMs are creating worlds that feel more like a sustained mood rather than a sequence of events, and that’s a significant shift in the narrative experience we are seeing.

Tom: It’s fascinating how the models prioritize ambiance over concrete action, and it really shows us where this bias is rooted in the data.

Jane: We've seen that specific pattern, but now let's look at how they measure this bias across time and space as we move into our next segment.

Improvements: Tom: Now, we’ve established that LLMs tend to overproduce perceived space and underproduce action space compared to human authors. The question is, what does the paper suggest we can do with this knowledge?

Jane: One key area of study is why this happens, which the analysis helps us understand by looking at things like prompt sensitivity. We find that while prompt phrasing matters, the overall spatial trends are quite robust across various versions of the instruction set.

Lu: That's a confidence boost for our data analysis, but it also highlights a challenge; even if we tweak the prompts, we’ aren't fixing the fundamental tendency of producing atmosphere over action. We still have to address that core bias in how they build worlds.

Meng: The authors suggest looking at "space-conditioned generation," which is a way to make space more intentional during prompting. This could potentially allow for much more controlled and deliberate worldbuilding, moving beyond the existing patterns of reliance on just the training data.

Lalam: I think the paper’s most important suggestion is that we need to guide the AI not just to learn from a baseline, but also to understand human preference for embodied interaction, which is what action space really represents.

Tom: That’s a huge idea, Lalam—forcing the AI to act like a novelist by requiring it to utilize action space in its generation is a powerful constraint.

Jane: We've covered the findings and the practical solutions, so let's bring all these insights together in our final summary of this paper.

Conclusion: Tom: So, wrapping up our discussion on "How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling," we have a very clear picture of the systematic differences between human and machine fiction. We’ve seen that LLMs strongly overrepresent perceived space while underproducing action space.

Jane: It really shows that these models aren't just mimicking style; they are constructing story worlds in a way that is fundamentally different from literary norms, which is a huge insight into the capabilities of AI.

Lu: The fact that this pattern holds across models and languages means we’ve found a consistent stylistic marker for AI-generated fiction, which opens up massive possibilities for understanding authorship.

Meng: And it gives us a benchmark—a way to measure how far our generative systems are from achieving that human level of grounded action and narrative momentum.

Lalam: We must remember the full implications of this paper and use its framework to guide the development so that AI can eventually achieve both atmospheric resonance and physical grounding.

Tom: It’s a really rich area of research, especially when we look at how these spatial metrics offer a path toward richer evaluation of AI-generated narratives.

Jane: It’s definitely going to be interesting to see what the next study on this topic reveals. Thank you all for joining us today and with that the final thoughts on "How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling."

Katrin Rohrbacher, Björn Nieth, Emmanuelle Salin, Bjoern Eskofier, Michaela Mahlberg

The Text and Language Lab, Department of Digital Humanities and Social Studies, Friedrich-Alexander University (FAU) Erlangen-Nürnberg · Department of Artificial Intelligence in Biomedical Engineering (AIBE), Friedrich-Alexander University (FAU) Erlangen-Nürnberg · Chair of AI-supported Therapy Decisions, Ludwig Maximilian University (LMU) Munich · Department of Linguistics and Communication, University of Birmingham · Munich Center for Machine Learning (MCML) · Institute of AI for Health, Helmholtz Centre Munich

cs.CL

Submitted: 2026-09-02

Updated: 2026-09-02

Code: https://github.com/BjoernNieth/worldbuilding-AI

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 86/100

The gist: The paper investigates the complex mechanisms by which Large Language Models (LLMs) construct and manage fictional worlds, focusing specifically on the deployment of setting and narrative space

Key concepts

Perceived Space
This is the tendency of LLMs to prioritize atmosphere or mood in their generated text. The study found that LLMs significantly overproduce this element compared to human writers, often reaching a normalized frequency twice the human baseline.
Action Space
This refers to the physical interaction and movement within a story's setting, which human authors typically ground their characters in. The analysis showed LLM-generated texts underproduce this concrete action space compared to human benchmarks.
Space-Conditioned Generation
This is a suggested method for prompting AI to make spatial elements more intentional during the generation process. It aims to allow for more controlled and deliberate worldbuilding beyond relying solely on existing training data.

Terminology

Summary

The paper investigates the complex mechanisms by which Large Language Models (LLMs) construct and manage fictional worlds, focusing specifically on the deployment of setting and narrative space across creative storytelling. By comparing AI-generated content with human-authored texts in both English and German, the research aims to determine if modern LLMs possess a deep structural understanding of spatial narrative mechanics or if their outputs merely reflect statistical pattern matching. This analysis is critical for advancing AI capabilities in creative fields, providing empirical evidence regarding the cognitive depth of generative models.

Spatial Category Distribution Across Narrative Sections

The study analyzes how setting types are distributed across different sections within a story (Figure 8). For human-authored fiction, the distribution provides a baseline frequency of spatial categories. The performance of various LLMs—including GPT 4.1, LlaMA 3.3, Mistral 3.2, and Gemma 3—is evaluated based on their normalized frequency of using categories such as action space, perceived space, and visual space. Furthermore, the models' ability to be identified based solely on their spatial usage is quantified using row-normalized confusion matrices (Figure 10). These matrices show that while models exhibit varying levels of overlap with human patterns, certain combinations of spatial category distributions allow for high model identification accuracy.

Temporal and Structural Consistency of Setting

The temporal distribution of setting categories is examined across narrative time for both English and German corpora (Figure 9). For instance, in the English corpus, the models show distinct patterns when comparing their usage across quarters (Q1–Q4) relative to the human baseline. Similar findings are observed in German, where models like GPT 4.1 and LlaMA 3.3 exhibit varying normalized frequencies of setting categories across the four quarters. Beyond temporal distribution, Figure 11 confirms that in human-authored corpora, Distributions are broadly stable across the main period of each corpus, suggesting a consistent historical use of setting types.

Perceived Space by Position within Chapter

The usage and placement of perceived space are analyzed quantitatively across chapter segments. Table 12 provides the normalized frequency of perceived space by quarter (Q1–Q4) for both English and German texts, averaged over four chapters. For example, in English, GPT 4.1 reports a mean perceived space of 8,456 words with a standard deviation (SD) of 324. When examining the position within the chapter (Figure 12), models are compared to the human baseline dashed line. The normalized frequency data for both English and German show that while models attempt to mimic human patterns, their specific distributions of perceived space across sections 1 through 4 vary significantly from the established human norm.

Overall Performance and Story Length

The study also provides general metrics on model output characteristics. Table 13 reports the story length in words for each model and language (n = 1,000 each). This data allows researchers to compare the overall scale of generated narratives across different architectures. Collectively, these findings demonstrate that while LLMs can generate text with complex spatial and temporal structures, their adherence to established human narrative patterns—particularly regarding the precise positioning and frequency of perceived space—remains a key area for further research.

Improvements for AI systems

Improvement: Implement a dedicated Structural Attention Module trained specifically on normalized frequency distributions across segmented text units (e.g., chapters or quarter-sections). This module must move beyond simple token prediction and model the probability distribution of specific discourse categories (spatial, temporal, action) relative to defined macro-structural boundaries.

What the improved AI system can do:

  • Genre Consistency Check: Accurately predict the expected normalized frequency of setting types (e.g., perceived space vs. visual space) for a given narrative segment (e.g., Chapter 2, Section 3) and flag deviations from established human-authored genre patterns (e.g., if a mystery novel suddenly deviates from the expected spatial distribution modeled by the human baseline in Figure 12).

  • Segmented Planning: When generating text, it can plan not just sentence-by-sentence, but section-by-section, ensuring that the required mix of descriptive, action, and perceived space types is met proportionally to maintain structural coherence across entire chapters or acts.


Abstract

In this paper, we analyze how Large Language Models (LLMs) employ worldbuilding strategies, focusing on setting as one measurable dimension of storyworld construction. We compare 1,000 AI-generated stories per model in English and German with human-authored fiction from Project Gutenberg. Building on prior work, we operationalize setting through five types of narrative space: "action", "perceived," "visual," "descriptive" and "no space", identified using fine-tuned BERT classifiers for German and English. We generate narratives using GPT 4.1, LlaMA 3.3, Mistral 3.2, and Gemma 3 and compare their spatial distributions to a human-authored baseline. We find that human-authored texts predominantly employ "action space," grounding narratives in embodied character-environment interaction, whereas LLMs systematically overproduce "perceived space," emphasizing atmosphere and affect. This divergence remains stable across narrative time. Overall, our findings show that LLMs exhibit worldbuilding patterns that differ consistently from human-authored fiction in ways that are both model-specific and language-sensitive.

Sources

Related papers