Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling".
Tom: LLMs have so far failed both to generate consistently compelling stories and to recognize this failure, which necessitates introducing narrative tension as a key metric for evaluating human-written fiction.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, to recap, this paper, "Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling," is essentially pointing out that LLMs struggle with generating stories that actually keep the reader guessing until the very end.
Jane: Exactly; the authors argue that current rubrics focus too much on visible drama rather than testing if information is actually being held back throughout the entire narrative, and they propose a new way to measure this using story endings.
Lu: The authors are looking at how we can make LLMs better at recognizing their own lack of tension, which is a key issue because model judges often just focus on things like urgent language instead of the underlying structure.
Meng: If the paper successfully links these narrative forecasting predictions to quality, it suggests that we need to start building evaluation benchmarks that look at the structural flow rather than just the final output or prose style.
Lalam: It’s about moving beyond simple factual checks; they are suggesting we need a metric that captures the feeling of uncertainty a reader experiences while reading fiction, which is something current systems miss entirely.
The paper's summary: Tom: The core summary of this research focuses on the one hundred-Endings metric, which tells us how often a model’s prediction of where a story will end matches the actual ending we have for that story <ref:2604.09854#pg0,the 100-Endings metric, which>.
Jane: They explain that tension is tied to predictability; if a story lacks tension, it becomes easy to guess what happens next, and this metric flags stories where the model’s predictions consistently fail to match reality.
Lu: They also introduce several complementary statistics alongside the main mismatch rate, like an inflection rate which tracks how often the prediction curve reverses direction, which directly relates to narrative manipulation or unexpected turns for the reader.
Meng: I'm interested in those secondary statistics because they might give us more diagnostic power than just a single number; knowing *why* a story is tense or not tense is much more useful for debugging.
Lalam: It really highlights that human readers don't read like checklists; they are immersed, and this paper tries to translate that holistic feeling of tension into something quantifiable through these predictive tests.
The paper's improvements: Tom: The authors propose using the one hundred-Endings metric as a way to evaluate quality because it directly measures narrative tension, which is what they believe is missing from most current rubrics <ref:2604.09854#pg0>.
Jane: They suggest that instead of relying on surface markers like urgent language, we should be rewarding the ability of a story to maintain uncertainty across its progression rather than just showing dramatic moments.
Lu: The authors demonstrate this by showing that when you apply a three-step generation pipeline grounded in literary theory—a warmup, an adaptation to the story idea, and then generation—tension significantly increases when measured by these forecasting methods.
Meng: That pipeline approach seems practical because it takes a vague idea and forces the model to think about structural elements like withholding information or escalating stakes before it even starts writing the final text.
Lalam: It shows that explicit, narratology-driven constraints can actually help LLMs achieve better long-range coherence, which is something we struggle with when we just let them write freely.
Conclusion: Tom: So, to wrap up, the main implication of this work is that we need to shift our focus in AI evaluation from just checking for prose quality to actively measuring narrative tension using tools like the one hundred-Endings metric <ref:2604.09854#pg0>.
Jane: They conclude that explicit, theory-driven work is necessary to map out and implement all the structural elements needed for what we might call true narrative machines, moving beyond simple pattern matching.
Lu: I think this opens up so many creative avenues for how we can guide story generation by incorporating these structural rules derived from literary theory directly into the prompt or training process.
Meng: From an engineering viewpoint, implementing a pipeline that adapts beats based on tension mechanisms could lead to much more robust and predictable long-form content generation in the future.
Lalam: It gives us a real roadmap for how to make our AI not just competent writers, but actually compelling storytellers by focusing on that core element of suspense.
McGill University · University of Chicago · University of British Columbia (Note: UBC is implied by the author list, though not explicitly stated as an organization in the header, but listed as an affiliation]
cs.CL
Submitted: 2026-04-10
Updated: 2026-10-03
Comments: COLM2026 Camera Ready. 44 pages, 10 figures, 30 tables
Code: https://github.com/EQ-bench/creative-writing-bench
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: LLMs have so far failed both to generate consistently compelling stories and to recognize this failure, which necessitates introducing narrative tension as a key metric for evaluating human-written
Key concepts
- 100-Endings Metric
- This metric measures narrative tension by testing a language model's ability to predict a story's ending 100 times after every sentence. Tension is defined by how often the model's predictions fail to match the actual, true ending of the story. High failure rates indicate high narrative unpredictability and tension.
- Narrative Recalcitrance
- This refers to a story's ability to resist easy prediction. When a story has strong narrative recalcitrance, it means that even with partial text, the ending is not easily anticipated by models. This property is considered the operational definition of genuine narrative tension.
- Pipeline for Improving Tension
- This is a three-step process designed to force LLMs to create tension. It starts by extracting structural beats from a reference story, adapts these beats to the target idea with specific instructions on withholding information and escalating stakes, and finally generates the story based on this detailed plan.
- Inflection Rate
- This statistic tracks narrative manipulation by measuring how often a story's tension curve sharply reverses direction at specific points. A high inflection rate shows that the story frequently surprises the reader or model by changing expectations at those moments.
Terminology
Summary
LLMs have so far failed both to generate consistently compelling stories and to recognize this failure, which necessitates introducing narrative tension as a key metric for evaluating human-written fiction. The core finding is that existing rubrics overlook tension, and a novel 100-Endings
metric operationalizes tension as unpredictability of story endings at the sentence level, revealing that professional fiction sustains narrative tension far better than LLM outputs.
The gist
Narrative tension is operationalized as unpredictability: after every sentence of a story, a model predicts how the story will end 100 times given only the text so far, and we measure tension as how often predictions fail to match the ground truth.
The Problem with Existing Evaluation
Current rubrics struggle to capture narrative tension because they reward surface markers of drama rather than verifying that information is genuinely withheld across the narrative arc. Rubric-based evaluation assumes story quality is an average over item scores, but human readers experience fiction holistically, feeling uncertain and compelled. Most rubrics omit tension as a criterion, yet model judges tend to focus on other aspects of writing. This gap means existing benchmarks fail to distinguish great stories from ones that just have good prose.
The 100-Endings Metric
The 100-Endings metric measures how predictable a story’s ending is at each point in its progression by prompting a language model to predict 100 possible endings and comparing those predictions against the held-out ground truth ending. The core intuition is that tension is operationally linked to predictability: a story that lacks tension will be easy to anticipate even from partial text.
This metric flags predictable stories when the model successfully anticipates their endings, while stories with genuine tension exhibit narrative recalcitrance that resists easy predictions.
Pipeline for Improving Tension
To address the gap, a generation pipeline grounded in literary theory is designed to inject tension. This pipeline consists of three sequential steps:
-
Step 0 — Warmup: Randomly select a reference story and extract
7 scene-level structural beats, each annotated with (1) the concrete action, (2) the specific narrative technique employed, and (3) what made it memorable.
-
Step 1 — Beat Adaptation: Adapt the extracted beat sheet to the target story idea to produce a
thick intermediate plan
specifyingthe narrative function, what information is revealed or withheld, the tension mechanism in play, and how stakes escalate.
-
Step 2 — Story Generation: Generate the final story conditioned on the structural to-do list from Step 1.
Experimental Results and Conclusion
The structured pipeline significantly increases narrative tension as measured by the 100-Endings metric, closing roughly half the gap to professional New Yorker fiction while maintaining performance on EQ-Bench. For instance, for Sonnet 4.6, the pipeline improves mean no-rate from 0.606 to 0.747 and late-stage no-rate more than doubles from 0.139 to 0.340 across three frontier models compared to vanilla generation (Table 2). The case study illustrates this: the zero-shot version of a romance story collapses tension early at 64% progression, whereas the pipeline version sustains high unpredictability through the final act, showing that the zero-shot failure is not an issue of prose fluency (both LLM versions are fluent), but a fundamental inability to defer closure.
The work concludes that explicit, narratology-driven work will be needed to map out and implement the full space of key narrative elements needed for true narrative machines.
Key Tension Statistics
The metric provides several complementary statistics beyond the mean no-rate:
- Inflection rate:
This is the fraction of evaluated positions at which the smoothed no-rate curve sharply reverses direction,
tracking narrative manipulation—the frequency with which a story reverses reader expectations.
- Post-spike retention:
For each local peak in the no-rate curve, this measures how much of the peak value is retained at the minimum within the next 10 positions.
Higher retention means tension peaks are sustained rather than immediately resolved.
The results show that New Yorker stories lead on all four metrics, with professional fiction sustaining unpredictability through the final act, as evidenced by a late-stage no-rate of 0.607 compared to 0.215 for Top-10 LLMs. The pipeline effect is consistent across models, demonstrating that literary theory-driven constraints can significantly improve long-range tension.
- Measurement Stability:
Stability tests confirm the metric's reliability, showing that mean no-rate and late-stage no-rate are stable to ±0.001 across four independent repeats of the generation process, indicating negligible measurement noise.
The pipeline effect is consistently observed across these repeats.
Improvements for AI systems
Here are specific improvements to current AI systems, derived directly from the methodology and findings of this paper:
- Improve LLM Story Generation for Narrative Tension (Targeting Creation)
The improved system will incorporate a three-stage, theory-grounded generation pipeline (Warmup Analysis → Beat Adaptation → Story Writing) to systematically increase narrative tension.
This system can do the following:
-
Identify and enforce structural elements like
information withholding
andtension mechanics
in the intermediate plan (Step 1). -
Generate stories that resist premature resolution, as demonstrated by the pipeline's ability to sustain high no-rates (e.g., achieving a late-stage no-rate of 0.747 for Sonnet 4.6 vs. its vanilla output).
-
Produce narratives where suspense is managed through subtext and delayed revelation rather than explicit, surface-level drama (as seen in the case study where the pipeline sustains tension between attraction without resolution).
- Develop a
Narrative Tension
Metric for Evaluation (Targeting Quality Assessment)
The system will replace or augment existing rubric-based judges with a position-level metric called the 100-Endings Metric.
This system can do the following:
-
Quantify narrative quality as the unpredictability of story endings at every sentence boundary.
-
Differentiate between stories that are merely well-written prose (high EQ-Bench scores) and stories that genuinely sustain tension (high no-rate).
-
Correct existing evaluation biases where judges favor resolution or surface drama by rewarding sustained narrative uncertainty, effectively creating a
gold standard
yardstick for literary fiction quality.
- Implement Structural Constraint Injection for Creativity (Targeting Idea Formulation)
The system will use the pipeline's ability to generate thick intermediate plans
derived from structural analysis of professional fiction (Step 1).
This system can do the following:
-
Transform vague story ideas into tension-rich stories by imposing a pre-defined narrative scaffolding that dictates pacing, conflict escalation, and information withholding.
-
Ensure that creative outputs are grounded in narratological principles (e.g., managing the fabula/syuzhet divide) rather than relying on stochastic generation alone.
- Enhance Model Alignment for Long-Form Coherence (Targeting Consistency)
By using the pipeline, LLMs can be explicitly trained to maintain long-range narrative coherence under structural constraints, addressing a known failure mode of current models.
This system can do the following:
- Significantly reduce plot collapse and premature closure in long narratives by forcing the model to adhere to a pre-planned structure across all beats.
- Create a
Tension Diagnostics
Tool (Targeting Debugging)
The system will allow researchers to diagnose why an LLM's story failed on tension metrics by analyzing its 100-Endings curve statistics (Mean No-Rate, Inflection Rate, Post-Spike Retention).
This tool can do the following:
-
Pinpoint whether a failure is due to a lack of initial plot uncertainty (low mean no-rate), rapid resolution (low post-spike retention), or narrative stagnation.
-
Provide specific feedback on which stage of the generation pipeline (Warmup, Adaptation, or Writing) contributed most to the observed tension deficit.
Abstract
LLMs have so far failed both to generate consistently compelling stories and to recognize this failure--on the leading creative-writing benchmark (EQ-Bench), LLM judges rank zero-shot AI stories above New Yorker short stories, a gold standard for literary fiction. We argue that existing rubrics overlook a key dimension of compelling human stories: narrative tension. We introduce the 100-Endings metric, which walks through a story sentence by sentence: at each position, a model predicts how the story will end 100 times given only the text so far, and we measure tension as how often predictions fail to match the ground truth. Beyond the mismatch rate, the sentence-level curve also yields complementary statistics that track plot-level twists and revelations, such as the inflection rate, a geometric measure of how frequently the curve reverses direction. Unlike rubric-based judges, 100-Endings correctly ranks New Yorker stories far above LLM outputs. Grounded in narratological principles, we design a story-generation pipeline using structural constraints, including analysis of story templates, idea formulation, and narrative scaffolding. Our pipeline significantly increases narrative tension as measured by the 100-Endings metric, while maintaining performance on the EQ-Bench leaderboard.
Sources
- Learning to Reason for Long-Form Story Generation
- Do Language Models Agree with Human Perceptions of Suspense in Stories?
- EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
- Frankentext: Stitching random text fragments into long-form narratives
- LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers
- SCORE: Story Coherence and Retrieval Enhancement for AI Narratives
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering