LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers".
Jane: LLMs consistently generate story continuations with 2–4× lower entropy and substantially higher context-dependence than human-authored ground truth—a gap that widens under post-training alignment and persists across model families and scales.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We just touched on the basic setup, but let's really unpack what this title means for us. It’s about LLMs exhibiting lower uncertainty in creative writing compared to professional writers who we know value ambiguity.
Jane: Exactly, Tom. The paper is pointing out a real disparity between how these systems generate text and how human artists actually create something rich.
Lu: It suggests that the way we align models today, which focuses on reducing uncertainty for factual accuracy, might be unintentionally steering us away from the very creative potential of literature.
Meng: I'm wondering if this means that when we ask an AI to write a story or a poem, it defaults to something that’s technically correct but lacks that human spark of unpredictability.
Lalam: I think what they mean is that the process of alignment itself is actively removing the very "openness" or multiple interpretations that make writing feel alive.
The paper's summary: Tom: So, if we look at the main points of this paper, it’s really about formalizing this "human-model uncertainty gap" using information theory to measure the difference between what humans write and what models produce.
Jane: That sounds complicated, Tom. Can you break down what that information-theoretic analysis actually tells us in plain English?
Lu: It quantifies the tension by looking at metrics like Mean Token Entropy and Perplexity, showing that human continuations are consistently more "surprising" to the model than the model's own generated text.
Meng: So, they’re saying that when you compare a human story continuation to an AI one under the same context, the human version is significantly more surprising for the AI to predict.
Lalam: It really highlights how instruction-tuned models struggle because they are optimized to minimize surprises and maximize predictability.
The paper's improvements: Tom: The paper suggests a few ways we can improve things, focusing on creating new uncertainty-aware alignment paradigms so the AI can actually capture that necessary ambiguity.
Jane: What kind of changes are they proposing for the way we train these models or fine-tune them? Are we talking about just tweaking some settings?
Lu: They propose implementing an Uncertainty-Aware Alignment Module, which would mean incorporating metrics like Token Entropy and PMI directly into the reinforcement learning from human feedback process.
Meng: From my side, that sounds like it would require retraining the reward model to actually value outputs that show high levels of controlled ambiguity rather than just low error rates.
Lalam: I think if we could train models to recognize when they need to introduce indeterminacy, like in a narrative tension building, instead of smoothing it out immediately, that would be a huge cultural improvement for AI writing.
Conclusion: Tom: So, wrapping up our discussion on "LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers," the paper concludes that achieving human-level creativity requires new alignment strategies that specifically account for this uncertainty gap.
Jane: It really brings us back to the idea that uncertainty isn't a flaw in writing; it’s a feature, and we need to adjust our systems to value it.
Lu: I think the main implication is that we need to move beyond simply aiming for factual correctness and start training models on how to handle conflicting meanings effectively.
Meng: For practical impact, this suggests that future creative AI tools shouldn't just be text generators but agents capable of understanding when to pause or introduce deliberate narrative friction.
Lalam: I think the biggest cultural implication is that if we can teach AI to embrace ambiguity as a necessary part of expression, it opens up entirely new ways for people to engage with digital art and storytelling.
cs.CL
Submitted: 2026-02-18
Updated: 2026-10-03
Importance score: 86/100
The gist: LLMs consistently generate story continuations with 2–4× lower entropy and substantially higher context-dependence than human-authored ground truth—a gap that widens under post-training
Key concepts
- Human–Model Uncertainty Gap
- This measures how much more surprising or unpredictable a human-written story is to an LLM compared to the LLM's own generated text. The paper found this gap is 2–4 times larger for human fiction, indicating models struggle to capture the full range of creative possibilities inherent in human writing.
- Information-Theoretic Framework
- This is a mathematical approach used to quantify uncertainty by analyzing information content. It uses metrics like Mean Token Entropy and Perplexity derived from log-probabilities to measure how much 'surprise' or predictive ambiguity exists in the text, allowing researchers to compare human and model outputs systematically.
- Divergence as Quality Driver
- The study found that higher divergence—meaning a continuation creates information distinct from the initial prompt—is positively correlated with quality scores. This implies that creative writing thrives when a text generates novel information rather than simply repeating strong, predictable patterns.
Terminology
Summary
LLMs consistently generate story continuations with 2–4× lower entropy and substantially higher context-dependence than human-authored ground truth—a gap that widens under post-training alignment and persists across model families and scales.
The Gist
LLMs consistently generate story continuations with 2–4× lower entropy and substantially higher context-dependence than human-authored ground truth—a gap that widens under post-training alignment and persists across model families and scales.
Theoretical Framework of Uncertainty in Creativity
Literary theory posits that uncertainty is a necessary condition for creative expression,
arising from concepts like indeterminacy, ambiguity, and the openness
of a work to multiple interpretations. This suggests that literary richness stems from the simultaneous presence of conflicting meanings and gaps in meaning-making. The paper formalizes this by proposing an information-theoretic framework to quantify the human–model uncertainty gap
between human-authored fiction and model-generated continuations. The core tension identified is that current alignment strategies, aiming to reduce ambiguity for instruction-following, steer models away from uncertain outputs to ensure factuality and reduce hallucination. This systematic aversion to uncertainty is described as a mechanism that effectively forecloses this co-creative process,
optimizing away a key precondition of the literary experience.
Experimental Methodology
The research employs a controlled story completion experiment comparing the uncertainty profiles of human-written versus model-generated sequences under identical settings. The comparison is structured by treating context as a cumulative prefix, constructing context–continuation pairs where the model's log-probability is evaluated for each token position in both the human continuation and the model continuation. To ensure a fair comparison, models are used as both generators and evaluators of human text uncertainty, with generation lengths constrained to match the ground truth length. Uncertainty is quantified using several metrics derived from log-probabilities:
-
Mean Token Entropy (Surprisal): Approximated using length-normalized Negative Log-Likelihood (NLL).
-
Perplexity (PPL): The exponentiated mean token surprisal, serving as a normalized measure of predictive uncertainty.
-
Pointwise Mutual Information (PMI): To disentangle context-dependent uncertainty from intrinsic token frequency, PMI normalizes conditional likelihood by an unconditional baseline.
-
Conditional PMI (CPMI): A thresholded variant that upweights the contribution of unconditional probabilities for tokens that are already uncertain given context.
Key Findings on Uncertainty Gaps
The analysis reveals a pervasive human–model uncertainty gap: when given identical contexts, LLMs consistently generate story continuations Tˆ2 with substantially lower intrinsic uncertainty than the human-authored ground truth T2.
This gap is quantified by NLL ratios ranging from 2.03 to 3.9 and PPL ratios amplifying this disparity to 2.76–8.82×, indicating that human-written stories are on average 2–4× more 'surprising' to the model than its own generations.
Furthermore, the gap is amplified in creative writing compared to other domains; for instance, the ∆PMI gap is 25–30% wider for creative writing (−2.15) compared to essays (−1.71) and news (−1.73).
The paper also notes that instruction-tuned and reasoning models exacerbate this trend compared to their base counterparts.
Correlation with Writing Quality
The study establishes a robust correlation between higher uncertainty metrics and automated quality scores, finding evidence of a possible “sweet spots” (inverted U-curve) where optimal writing quality exists at high entropy levels. For human ground truth, the mean Spearman correlation between token-level uncertainty (NLL, PPL) and quality scores is positive: for human-authored continuations, the mean Spearman correlation is ρ¯ = 0.072.
Crucially, PMI shows a robust negative correlation with quality for both human and model text (PMI shows a robust negative correlation with quality for both human (ρ¯ = −0.106, 75% significantly negative) and model text (ρ¯ = −0.076, 69% significantly negative)
). This suggests that divergence is a linear driver of quality: the more a continuation creates its own information distinct from the strong priors of the prompt, the higher it is rated.
The quadratic regression analysis confirms this non-monotonic relationship for human text, with optimal uncertainty occurring at z¯∗ = 1.98 standard deviations above the mean.
Domain Generalization and Model Effects
The uncertainty gap is not unique to creative writing but is amplified in creative writing, especially on context-dependent uncertainty metrics.
The analysis across three domains (Creative Writing, News, Essays) confirms that the ∆PMI gap is consistently larger for creative writing.
Improvements for AI systems
Here are specific improvements to AI systems based on the findings of this research:
-
Enhance creative generation capabilities by implementing a
Uncertainty-Aware Alignment Module.
This module should explicitly incorporate information-theoretic metrics (Token Entropy, Perplexity, PMI, and CPMI) during the reinforcement learning from human feedback (RLHF) or alignment phase. -
Implement a dynamic uncertainty thresholding mechanism in the decoding process for generative models. Instead of forcing outputs into a low-entropy
safe
regime immediately, the system should be capable of sampling based on a calculated uncertainty metric that allows for controlled bursts of high-entropy, creative exploration when the context demands ambiguity (e.g., during narrative tension building). -
Develop a quality assessment feedback loop that specifically weights outputs based on their
creative uncertainty profile.
The reward model should be fine-tuned to positively correlate with the observedsweet spot
in the uncertainty-quality relationship (the inverted-U curve), thereby rewarding outputs that exhibit optimal, controlled ambiguity rather than those that are perfectly predictable or maximally random. -
For LLMs used in narrative co-creation, develop a mechanism to distinguish between destructive hallucinations and constructive ambiguity. This involves training the model not just on factual correctness but also on
literary affordances
—learning when and how to introduce indeterminacy (unbestimmtheitsstellen) or ambiguity (Empson's gridiron) as a necessary aesthetic feature of the text. -
Integrate domain-specific uncertainty calibration into alignment strategies. Since the paper shows that the human–model uncertainty gap is amplified in creative writing compared to functional domains (News/Essays), future systems should utilize domain classifiers to apply more aggressive, uncertainty-aware constraints specifically for open-ended, narrative tasks while relaxing them for structured reasoning or factual retrieval tasks.
-
Improve model interpretability by providing
Uncertainty Attribution Maps.
When a model generates a text segment, it should be able to highlight which tokens contributed most significantly to the local surprisal (Htoken) and how contextual relevance (PMI/CPMI) was calculated for that specific decision point, allowing researchers to diagnose why the model chose an ambiguous or predictable path.
These improvements will enable AI systems to move beyond merely generating safe
or average
text, allowing them to produce creative writing that exhibits the necessary features of literary richness: tension, ambiguity, and genuine surprise.
Sources
- TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories
- Gemma 2: Improving Open Language Models at a Practical Size
- Learning to Reason for Long-Form Story Generation
- Creative Writing with an AI-Powered Writing Assistant: Perspectives from Professional Writers
- Evaluating Creative Short Story Generation in Humans and Large Language Models
- Why Language Models Hallucinate
- Rethinking Creativity Evaluation: A Critical Analysis of Existing Creativity Evaluations
- Creativity Has Left the Chat: The Price of Debiasing Language Models
- Mind the Gap: Conformative Decoding to Improve Output Diversity of Instruction-Tuned Large Language Models
- Frankentext: Stitching random text fragments into long-form narratives
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering