Novels generated by language models show compressed formal variation
Mehdy Sedaghat Payam, Justin Quinn
University of Maryland · University of West Bohemia · Charles University
cs.CL, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-14
Project page: https://libraryponders.github.io
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 75/100
The gist: This study investigates whether large language models (LLMs) can produce the same range of formal diversity across complete novels as human-written corpora, rather than focusing on detecting
Terminology
Summary
This study investigates whether large language models (LLMs) can produce the same range of formal diversity across complete novels as human-written corpora, rather than focusing on detecting individual AI-generated passages. The paper contrasts six corpora: twenty novels generated using GPT-5.5 Thinking in a nineteenth-century British realist style, twenty novels generated using Qwen3-14B in the same style, twenty novels generated using each model in a contemporary zero style,
205 nineteenth-century human-written British novels, and sixty-five contemporary human-written Zero-Style novels.
At the document level, the research measures MATTR-500 (local lexical diversity), Shannon entropy, average sentence length, readability (Flesch Reading Ease), and punctuation rate. The most robust and reliable result is compression of sentence structure: Repeated generations produce novels that vary far less from one another in sentence structure than human novels do.
Compression is also present in readability, punctuation, and sentence length variability within novels. Lexical measures tend to be similarly compressed, with the exception of Qwen Zero-Style MATTR.
The paper introduces the term overclosure
to describe a corpus-level condition where generated texts remain locally fluent but show an unusually restricted range across measured formal dimensions. It distinguishes between variance overclosure
(limited formal range between novels) and correlational overclosure
(features becoming more strongly coordinated in generated texts than in human comparators). The study finds strong evidence for variance overclosure but not for correlational overclosure, as cross-feature correlations vary across models and stylistic conditions without a stable pattern.
Key quantitative results include: of sixteen raw ratios of AI-generated to human standard deviation, fifteen are under 0.58, and fifteen Brown-Forsythe comparisons remain under q =.05 after adjustment. The one exception is Qwen Zero-Style MATTR (ratio 0.576, adjusted q =.073). Sentence-length SD ratios are extremely low: 0.059 for GPT Victorian, 0.059 for Qwen Victorian, 0.067 for GPT Zero-Style, and 0.116 for Qwen Zero-Style. Intra-novel sentence-length dispersion ratios are even lower (0.051, 0.034, 0.052, and 0.089 respectively). These results persist after author-balanced bootstrap resampling, length adjustment, leave-one-author-out sensitivity tests, and restricting the historical comparator to 1837–1900.
Despite compression, GPT and Qwen do not converge on a shared style. They exhibit distinct mean stylistic profiles: GPT has larger lexical means, Qwen has smaller entropy, Qwen is more readable, and punctuation patterns differ substantially. For example, GPT Victorian has higher MATTR and entropy than the historical human corpus, while Qwen Victorian is close to the human MATTR mean but has lower entropy. The paper emphasizes that overclosure is not about low metric values or literary quality—it is purely about reduced formal range.
The qualitative close reading of three AI-generated novels (Where the Bell Calls, The House Between Two Tongues, and The Quota Line) reveals that the statistical compression is not perceptible to readers; the prose often resembles recently unearthed classics of Victorian realism.
However, the close reading identifies recurring wit in dialogue and frequent use of syllepsis (zeugma) across the novels, which contributes to stylistic monotony, though this is unlikely to affect the measured metrics. The study also notes structural patterns such as marriage plots resolving early and episodic organization in later sections.
The paper concludes that variance overclosure is the principal finding, observed across both model–workflow configurations and both stylistic targets. Sentence-length compression is the most robust dimension, remaining unusually low even after all robustness checks. Correlational overclosure remains a conditional, exploratory hypothesis requiring further research with more model families and matched prompt sets. The findings are limited to the specific models, interfaces, prompts, topics, and generation processes considered, and should not be interpreted as a universal estimate of language-model fiction.
Improvements for AI systems
Improvements to AI systems:
-
Implement variance-aware decoding for sentence-length diversity. Modify the sampling or beam-search process to explicitly track and enforce a target distribution of sentence lengths across a generated document, preventing the observed compression (SD ratios as low as 0.034–0.116 vs. human). The improved system can generate novels with human-like intra- and inter-novel sentence-length variability.
-
Add a corpus-level
overclosure
regularizer during multi-document generation. When generating a batch of texts (e.g., a book series or a collection), monitor the standard deviation of MATTR, readability, and punctuation rate across outputs. If the ratio of AI-to-human SD falls below 0.58, inject controlled noise into stylistic parameters (e.g., vocabulary richness, punctuation frequency) for subsequent documents. The improved system can produce a set of novels with a formal range comparable to human corpora, not just locally fluent individual texts. -
Introduce a
stylistic drift
mechanism to prevent cross-novel monotony. Since GPT and Qwen show distinct mean profiles but low variance, the system can learn a per-modelstyle fingerprint
and then deliberately perturb it along principal components of human stylistic variation (e.g., from the 205-novel historical corpus). The improved system can generate multiple novels that are recognizably from the same model but vary in sentence structure, readability, and lexical diversity as much as different human authors do. -
Develop a
syllepsis-aware
stylistic diversity module. Because close reading revealed recurring wit and syllepsis as a source of monotony, the system can parse generated dialogue and narrative for repeated rhetorical figures (zeugma, syllepsis) and actively reduce their frequency or vary their syntactic placement. The improved system can avoid stylistic tics that are invisible to statistical metrics but noticeable to human readers. -
Create a
correlational overclosure
diagnostic tool. Since cross-feature correlations (e.g., sentence length vs. readability) vary unpredictably across models, the system can compute a correlation matrix of formal features after each generation and compare it to human baselines. If the matrix is more tightly coordinated (e.g., all features moving together), the system can re-randomize feature combinations. The improved system can detect and correct for hidden over-coordination that current single-metric evaluations miss. -
Enable
targeted historical-style variance
generation. Given that GPT Victorian had higher MATTR and entropy than human historical novels, the system can calibrate its output to match not just the mean but the distribution of a target era's formal features. The improved system can generate a Victorian-style novel that is statistically indistinguishable from a human-written one across all five measured dimensions, including their natural spread. -
Add a
post-hoc variance correction
pipeline. After generating a novel, the system can compute its sentence-length SD, readability SD, and punctuation SD against a human reference corpus. If any ratio falls below 0.5, it can rewrite specific sections (e.g., lengthen or shorten sentences, adjust punctuation density) to restore human-like dispersion without altering meaning. The improved system can output a final text that passes both local fluency and global diversity checks.
Abstract
While large language models can generate entire novels, there is little information about the level of formal variation in their output over many generations. Rather than asking whether individual passages can be identified as AI-generated, this study asks whether repeated AI generation can produce the same range of diversity which is found across human corpora. This paper contrasts six corpora based on generation source and target style: twenty novels generated using GPT-5.5 Thinking in a nineteenth-century British realist style, twenty novels generated using Qwen3-14B in a nineteenth-century British realist style, twenty novels generated using each of these models in a contemporary zero style, 205 nineteenth-century human-written British novels, and sixty-five contemporary human-written Zero-Style novels. At the document level, the research includes MATTR-500, Shannon entropy, average sentence length, readability, and punctuation rate measurements. The most robust and reliable result is compression of sentence structure. Repeated generations produce novels that vary far less from one another in sentence structure than human novels do. Compression is also present in the measures of readability, punctuation, and sentence length variability within novels. Lexical measures tend to be similarly compressed, with the exception of Qwen Zero-Style MATTR. Despite having distinct mean stylistic profiles, GPT and Qwen lack a stable pattern of cross-measure correlation. This article therefore distinguishes between variance overclosure, which represents a limited formal range between novels, and a more specific phenomenon of correlational overclosure. This means that an individual AI-generated novel may resemble human fiction stylistically, while a collection of AI-generated novels occupies a much narrower formal range.
Sources
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering