LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers
summary
The gist
LLMs consistently generate story continuations with 2–4× lower entropy and substantially higher context-dependence than human-authored ground truth—a gap that widens under post-training
In short
The research compared human-written stories with those generated by LLMs to measure uncertainty. Findings show LLM continuations have significantly lower intrinsic uncertainty than human text, a gap that widens in creative writing. This suggests current alignment methods inadvertently suppress the necessary ambiguity and 'openness' required for rich, creative expression.
Key concepts
- Human–Model Uncertainty Gap
- This measures how much more surprising or unpredictable a human-written story is to an LLM compared to the LLM's own generated text. The paper found this gap is 2–4 times larger for human fiction, indicating models struggle to capture the full range of creative possibilities inherent in human writing.
- Information-Theoretic Framework
- This is a mathematical approach used to quantify uncertainty by analyzing information content. It uses metrics like Mean Token Entropy and Perplexity derived from log-probabilities to measure how much 'surprise' or predictive ambiguity exists in the text, allowing researchers to compare human and model outputs systematically.
- Divergence as Quality Driver
- The study found that higher divergence—meaning a continuation creates information distinct from the initial prompt—is positively correlated with quality scores. This implies that creative writing thrives when a text generates novel information rather than simply repeating strong, predictable patterns.
Terminology used across episodes
This episode discusses
- LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers · Paper Radio
- TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories
- Gemma 2: Improving Open Language Models at a Practical Size
- Learning to Reason for Long-Form Story Generation
- Creative Writing with an AI-Powered Writing Assistant: Perspectives from Professional Writers
- Evaluating Creative Short Story Generation in Humans and Large Language Models
- Why Language Models Hallucinate
- Rethinking Creativity Evaluation: A Critical Analysis of Existing Creativity Evaluations
- Creativity Has Left the Chat: The Price of Debiasing Language Models
- Mind the Gap: Conformative Decoding to Improve Output Diversity of Instruction-Tuned Large Language Models
- Frankentext: Stitching random text fragments into long-form narratives
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
The paper
LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers · Read on arXiv
We argue that uncertainty is a key and understudied limitation of LLMs' performance in creative writing, which is often characterized as trite and cliché-ridden. Literary theory identifies uncertainty as a necessary condition for creative expression, while current alignment strategies steer models away from uncertain outputs to ensure factuality and reduce hallucination. We formalize this tension by quantifying the ``uncertainty gap'' between human-authored stories and model-generated continuations. Through a controlled information-theoretic analysis of 28 LLMs on high-quality storytelling datasets, we demonstrate that human writing consistently exhibits significantly higher uncertainty than model outputs. We find that instruction-tuned and reasoning models exacerbate this trend compared to their base counterparts; furthermore, the gap is more pronounced in creative writing than in functional domains, and shows a consistent correlation with writing quality. Achieving human-level creativity requires new uncertainty-aware alignment paradigms that can distinguish between destructive hallucinations and the constructive ambiguity required for literary richness.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers".
Jane: LLMs consistently generate story continuations with 2–4× lower entropy and substantially higher context-dependence than human-authored ground truth—a gap that widens under post-training alignment and persists across model families and scales.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We just touched on the basic setup, but let's really unpack what this title means for us. It’s about LLMs exhibiting lower uncertainty in creative writing compared to professional writers who we know value ambiguity.
Jane: Exactly, Tom. The paper is pointing out a real disparity between how these systems generate text and how human artists actually create something rich.
Lu: It suggests that the way we align models today, which focuses on reducing uncertainty for factual accuracy, might be unintentionally steering us away from the very creative potential of literature.
Meng: I'm wondering if this means that when we ask an AI to write a story or a poem, it defaults to something that’s technically correct but lacks that human spark of unpredictability.
Lalam: I think what they mean is that the process of alignment itself is actively removing the very "openness" or multiple interpretations that make writing feel alive.
The paper's summary: Tom: So, if we look at the main points of this paper, it’s really about formalizing this "human-model uncertainty gap" using information theory to measure the difference between what humans write and what models produce.
Jane: That sounds complicated, Tom. Can you break down what that information-theoretic analysis actually tells us in plain English?
Lu: It quantifies the tension by looking at metrics like Mean Token Entropy and Perplexity, showing that human continuations are consistently more "surprising" to the model than the model's own generated text.
Meng: So, they’re saying that when you compare a human story continuation to an AI one under the same context, the human version is significantly more surprising for the AI to predict.
Lalam: It really highlights how instruction-tuned models struggle because they are optimized to minimize surprises and maximize predictability.
The paper's improvements: Tom: The paper suggests a few ways we can improve things, focusing on creating new uncertainty-aware alignment paradigms so the AI can actually capture that necessary ambiguity.
Jane: What kind of changes are they proposing for the way we train these models or fine-tune them? Are we talking about just tweaking some settings?
Lu: They propose implementing an Uncertainty-Aware Alignment Module, which would mean incorporating metrics like Token Entropy and PMI directly into the reinforcement learning from human feedback process.
Meng: From my side, that sounds like it would require retraining the reward model to actually value outputs that show high levels of controlled ambiguity rather than just low error rates.
Lalam: I think if we could train models to recognize when they need to introduce indeterminacy, like in a narrative tension building, instead of smoothing it out immediately, that would be a huge cultural improvement for AI writing.
Conclusion: Tom: So, wrapping up our discussion on "LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers," the paper concludes that achieving human-level creativity requires new alignment strategies that specifically account for this uncertainty gap.
Jane: It really brings us back to the idea that uncertainty isn't a flaw in writing; it’s a feature, and we need to adjust our systems to value it.
Lu: I think the main implication is that we need to move beyond simply aiming for factual correctness and start training models on how to handle conflicting meanings effectively.
Meng: For practical impact, this suggests that future creative AI tools shouldn't just be text generators but agents capable of understanding when to pause or introduce deliberate narrative friction.
Lalam: I think the biggest cultural implication is that if we can teach AI to embrace ambiguity as a necessary part of expression, it opens up entirely new ways for people to engage with digital art and storytelling.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck