LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers

summary

Video file (mp4)

The gist

LLMs consistently generate story continuations with 2–4× lower entropy and substantially higher context-dependence than human-authored ground truth—a gap that widens under post-training

In short

The research compared human-written stories with those generated by LLMs to measure uncertainty. Findings show LLM continuations have significantly lower intrinsic uncertainty than human text, a gap that widens in creative writing. This suggests current alignment methods inadvertently suppress the necessary ambiguity and 'openness' required for rich, creative expression.

Key concepts

Human–Model Uncertainty Gap
This measures how much more surprising or unpredictable a human-written story is to an LLM compared to the LLM's own generated text. The paper found this gap is 2–4 times larger for human fiction, indicating models struggle to capture the full range of creative possibilities inherent in human writing.
Information-Theoretic Framework
This is a mathematical approach used to quantify uncertainty by analyzing information content. It uses metrics like Mean Token Entropy and Perplexity derived from log-probabilities to measure how much 'surprise' or predictive ambiguity exists in the text, allowing researchers to compare human and model outputs systematically.
Divergence as Quality Driver
The study found that higher divergence—meaning a continuation creates information distinct from the initial prompt—is positively correlated with quality scores. This implies that creative writing thrives when a text generates novel information rather than simply repeating strong, predictable patterns.

Terminology used across episodes

This episode discusses

The paper

LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers · Read on arXiv

We argue that uncertainty is a key and understudied limitation of LLMs' performance in creative writing, which is often characterized as trite and cliché-ridden. Literary theory identifies uncertainty as a necessary condition for creative expression, while current alignment strategies steer models away from uncertain outputs to ensure factuality and reduce hallucination. We formalize this tension by quantifying the ``uncertainty gap'' between human-authored stories and model-generated continuations. Through a controlled information-theoretic analysis of 28 LLMs on high-quality storytelling datasets, we demonstrate that human writing consistently exhibits significantly higher uncertainty than model outputs. We find that instruction-tuned and reasoning models exacerbate this trend compared to their base counterparts; furthermore, the gap is more pronounced in creative writing than in functional domains, and shows a consistent correlation with writing quality. Achieving human-level creativity requires new uncertainty-aware alignment paradigms that can distinguish between destructive hallucinations and the constructive ambiguity required for literary richness.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers".

Jane: LLMs consistently generate story continuations with 2–4× lower entropy and substantially higher context-dependence than human-authored ground truth—a gap that widens under post-training alignment and persists across model families and scales.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We just touched on the basic setup, but let's really unpack what this title means for us. It’s about LLMs exhibiting lower uncertainty in creative writing compared to professional writers who we know value ambiguity.

Jane: Exactly, Tom. The paper is pointing out a real disparity between how these systems generate text and how human artists actually create something rich.

Lu: It suggests that the way we align models today, which focuses on reducing uncertainty for factual accuracy, might be unintentionally steering us away from the very creative potential of literature.

Meng: I'm wondering if this means that when we ask an AI to write a story or a poem, it defaults to something that’s technically correct but lacks that human spark of unpredictability.

Lalam: I think what they mean is that the process of alignment itself is actively removing the very "openness" or multiple interpretations that make writing feel alive.

The paper's summary: Tom: So, if we look at the main points of this paper, it’s really about formalizing this "human-model uncertainty gap" using information theory to measure the difference between what humans write and what models produce.

Jane: That sounds complicated, Tom. Can you break down what that information-theoretic analysis actually tells us in plain English?

Lu: It quantifies the tension by looking at metrics like Mean Token Entropy and Perplexity, showing that human continuations are consistently more "surprising" to the model than the model's own generated text.

Meng: So, they’re saying that when you compare a human story continuation to an AI one under the same context, the human version is significantly more surprising for the AI to predict.

Lalam: It really highlights how instruction-tuned models struggle because they are optimized to minimize surprises and maximize predictability.

The paper's improvements: Tom: The paper suggests a few ways we can improve things, focusing on creating new uncertainty-aware alignment paradigms so the AI can actually capture that necessary ambiguity.

Jane: What kind of changes are they proposing for the way we train these models or fine-tune them? Are we talking about just tweaking some settings?

Lu: They propose implementing an Uncertainty-Aware Alignment Module, which would mean incorporating metrics like Token Entropy and PMI directly into the reinforcement learning from human feedback process.

Meng: From my side, that sounds like it would require retraining the reward model to actually value outputs that show high levels of controlled ambiguity rather than just low error rates.

Lalam: I think if we could train models to recognize when they need to introduce indeterminacy, like in a narrative tension building, instead of smoothing it out immediately, that would be a huge cultural improvement for AI writing.

Conclusion: Tom: So, wrapping up our discussion on "LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers," the paper concludes that achieving human-level creativity requires new alignment strategies that specifically account for this uncertainty gap.

Jane: It really brings us back to the idea that uncertainty isn't a flaw in writing; it’s a feature, and we need to adjust our systems to value it.

Lu: I think the main implication is that we need to move beyond simply aiming for factual correctness and start training models on how to handle conflicting meanings effectively.

Meng: For practical impact, this suggests that future creative AI tools shouldn't just be text generators but agents capable of understanding when to pause or introduce deliberate narrative friction.

Lalam: I think the biggest cultural implication is that if we can teach AI to embrace ambiguity as a necessary part of expression, it opens up entirely new ways for people to engage with digital art and storytelling.

More episodes

← Home