ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control

arXiv:2608.05169 · cs.CL · Submitted 2026-05-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control".

Jane: The paper was written by Jindong Lia, Yang Yang, Zihao Liu, Yutao Yue and Menglin Yangb* from Hong Kong University of Science and Technology (GZ).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary and Implications: Jane: So, the authors’ summary of ConWriter is that it treats long-form writing not as a single freeform generation process, but as an incremental state-transition task. It's a fundamental change in how we view AI authorship.

Tom: Exactly, Jane; they are guiding the LLM through a highly structured process instead of just letting it run wild. The paper explains that this framework uses two distinct memories—a static memory for things like world rules, and dynamic memory for everything currently happening in the story.

Lu: This dual-memory model is key to maintaining long-range coherence. It allows the system to separate immutable facts, like "this castle exists," from evolving state changes, like "the characters have moved past this point."

Meng: I’m curious about how this state transition is actually checked. It sounds like they are deriving a symbolic operator for every scene that must be satisfied before accepting the output. It's a rigorous validation process.

Lalam: This mechanism ensures that the narrative isn't just rambling; it's moving forward correctly, following a path dictated by its own internal rules and commitments, which is deeply satisfying from a storytelling perspective.

Jane: It’s truly an active form of memory; you aren're giving the AI not just a context window, but an evolving set of governing laws for the entire narrative world. This capability is going to revolutionize how we approach complex storytelling.

Tom: And this brings us to discussing the hard data—how does this system actually perform compared to the results they show in their experiments?

Improvements and Findings: Jane: The quantitative results are incredibly compelling, Tom. They show that ConWriter consistently reduces the overall consistency error density, or CED, across all tested base models. It's not just a marginal improvement; it’s a substantial leap in reliability.

Tom: It’s clear that this framework is designed to handle the challenges of long-form generation where things like timeline and characterization errors tend to accumulate over thousands of words. The paper really shows how it prevents that drift from happening.

Lu: I think what the numbers demonstrate, however, is more than just reducing errors; it’s proving that the complexity of state management is effective at mitigating cumulative drift over long-range dependencies. It’s about quantifying reliability over distance.

Meng: I was particularly interested in a finding mentioned regarding the forced-length setting—that while we mandate a specific word count, ConWriter sometimes generates more consistent content than requested. This implies an inherent robustness in the planning phase.

Lalam: That suggests that the system isn't just filling slots to meet a quota; it’s organically expanding the narrative in a way that is perfectly aligned with its own internal rules, which is a massive indicator of quality control and artistic depth.

Jane: It really confirms that this architecture acts as a robust consistency layer, applicable even if the base model—like GPT-five point four-nano—is already quite reliable. This suggests broad utility across various AI tools and deployment scenarios.

Tom: So, we’ve seen the architecture and the results. Now, let's get into the nitty-gritty of how they tested this and what specific insights those experiments provide.

Detailed Experimental Analysis: Jane: The experiments cover four different tasks—continuation, generation, expansion, and completion—and three target lengths from 3K up to 12K words. This multi-length setting is vital for testing long-range consistency.

Tom: And the data shows that ConWriter’s performance is particularly strong when the base model exhibits clear consistency errors. For example, with DeepSeek-V4-Flash, the reduction in CED was massive—over eighty-seven percent at 3K words!

Lu: I think we need to look at their ablation studies here. The fact that removing dynamic memory causes a huge spike in error density proves conclusively that state tracking is not a luxury; it's absolutely essential for narrative integrity.

Meng: It’s interesting because the gains are so dramatic when the base model is struggling, but Lalam, what does this tell us about where this technology fits into our actual workflow?

Lalam: This gives us a tool that allows us to handle complex plots with absolute structural integrity. It means we can build stories that feel naturally evolved and deeply consistent, allowing creators to focus on the art rather than the mechanics of maintaining consistency.

Jane: Exactly, Lalam; the system is designed to be a crucial layer of reliability. The combination of symbolic checking and risk monitoring ensures that even if local errors are detected, we have a way to fix them before they become catastrophic failures.

Tom: And this leads us perfectly into our final wrap-up as we summarize the overall impact of all these findings today.

Conclusion: Jane: To summarize, "ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control" offers a powerful solution to narrative drift. It's a reliable, state-aware approach for any complex long-form content creation.

Tom: Absolutely, Jane; it sets a new standard for what we expect from AI authorship—moving beyond just having the ability to generate text to having the capability to reliably manage that evolving state.

Lu: I’m just thrilled that we’re seeing a way to blend rigorous symbolic logic with AI creativity, transforming what used to be pure luck in long-form writing into something highly engineered and predictable.

Meng: It’s encouraging for practical application; it makes these massive, long-form projects much more feasible and manageable because the technical constraints are clearly defined.

Lalam: Our final thought is that ConWriter gives us a tool that helps us craft stories with both deep complexity and absolute structural integrity, allowing creators to achieve high narrative goals.

Tom: It sounds like this is a monumental step forward in AI generation, so thank you all for sharing your insights on "ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control."

Jane: I'm looking forward to discussing the next breakthrough in AI with all of you soon.

Hong Kong University of Science and Technology (GZ)

cs.CL

Submitted: 2026-05-27

Updated: 2026-09-04

Comments: Accepted to Findings of EMNLP 2026

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 85/100

The gist: " * Abstract and Problem Formulation Long stories inherently suffer from accumulated inconsistency, where existing prompting-based methods often fail to maintain coherence.

Key concepts

ConWriter
ConWriter is a framework that treats long-form writing not as freeform generation, but as an incremental state-transition task. It guides the LLM through a highly structured process to ensure the narrative moves forward correctly and maintains internal coherence.
Dual Memory Model
This model uses two distinct memories: static memory stores immutable facts or world rules (e.g., 'this castle exists'), while dynamic memory tracks all current, evolving state changes within the the story.
Consistency Error Density (CED)
CED is a quantitative metric used to measure the reliability of AI generation. ConWriter was designed to substantially reduce this density across various models, proving its effectiveness in mitigating errors over long-range dependencies.
Neuro-Symbolic Control
This refers to the rigorous validation process where symbolic logic checks every scene before accepting output. This ensures the narrative follows a path dictated by its own internal rules, providing structural integrity to the complex story.

Terminology

Summary

"


Abstract and Problem Formulation

Long stories inherently suffer from accumulated inconsistency, where existing prompting-based methods often fail to maintain coherence. The paper identifies that Long stories accumulate inconsistency over time and that "long-form story generation requires models to preserve narrative consistency across extended contexts, yet existing prompting-based methods often accumulate temporal, factual, character, commonsense, and stylistic errors as the story grows. The core challenge is that a local error often triggers large-scale regeneration," necessitating a mechanism to control consistency during generation.

Proposed Solution: ConWriter

The authors propose CONWRITER, which they define as a training-free framework for consistency-aware long story generation. CONWRITER addresses the problem by treating long-story writing not as a single freeform decoding process, but as an incremental, stateful process.

Core Mechanism and Methodology

CONWRITER operates by writing the story scene by scene, guided by four lightweight control mechanisms:

  1. Dual-Memory Modeling: This mechanism separates immutable story commitments from evolving events, entity states, relations, timelines, and pending constraints. It maintains a combination of static memory (M s) and dynamic memory (M dt).

  2. Symbolic State-Transition Reasoning: Each scene is represented as a checkable transition over narrative states. The framework derives a symbolic transition operator (o t = (Pre t, Post t, Forbid t)) which specifies the conditions that must hold before and after the scene.

  3. Dual Consistency Assurance: This combines structured validation with uncertainty-aware risk monitoring. The system uses this assurance to inspect both explicit violations and weakly grounded segments.

  4. Sentence-Level Patch Repair: Instead of regenerating entire scenes, CONWRITER utilizes this mechanism to revise only conflict-bearing sentences before accepted scenes are committed into dynamic memory, using symbolic violation anchors and entropy-based risk signals.

The Generation Loop (Process)

The generation process is formalized as a constrained incremental writing loop. At each step, the framework:

  1. Retrieves relevant memory (R t).

  2. Reasons to derive the symbolic transition operator (o t).

  3. Generates a draft scene (t) using the base LLM (G theta), guided by memory and constraints.

  4. The draft is examined by dual consistency assurance, which checks if the scene realizes a valid narrative transition.

  5. If violations are found, the local repair module is activated to apply a minimal patch (Patch only conflict-bearing sentences).

  6. The scene (y t) is committed to dynamic memory only after it passes consistency checks, ensuring that local inconsistencies from propagating into later scenes.

Evaluation and Results

The framework was evaluated on ConStory-Bench (Li et al., 2026), covering four long-story tasks: continuation, generation, expansion, and completion. The experiments were conducted across three base LLM families (Qwen3.5-Plus, DeepSeek-V4-Flash, and GPT-5.4-nano) at target lengths of 3K, 6K, and 12K words.

The primary findings indicate that:

  • C ON W RITER consistently reduces Overall Avg. CED across the three base LLM families.

  • The gains are particularly pronounced when the base model exhibits clear consistency errors (e.g, DeepSeek-V4-Flash).

  • The results show that the improvement is not merely caused by generating shorter or simpler stories, as CONWRITER can maintain lower consistency error density while supporting extended story development.

Ablation and Discussion

A detailed ablation study confirms the necessity of all components: removing dynamic memory, structured validation, or uncertainty monitoring leads to a significant degradation in consistency (higher Avg. CED). The authors conclude that C ON W RITER mitigates this issue by validating each scene before committing it into dynamic memory, rather than allowing local errors to propagate into later scenes, providing a robust solution for long-form narrative coherence.

Improvements for AI systems

The following improvements are derived from the principles and mechanisms of ConWriter (Transition-Constrained State-ful Long-Form Story Generation) and represent targeted enhancements applicable to any AI system tasked with long-form content generation.


Improvement: Instead of relying solely on a large, monolithic context window (which suffers from memory decay), the system must rigidly separate narrative data into two distinct, structured memory banks:

  • Static Memory (M s): Stores immutable constraints (e.g., world rules, character core attributes, defined style/tone commitments). This acts as a fixed ground truth against drift.

  • Dynamic Memory (M dt): Stores the evolving state (e.g., current location of entities, recent actions, unresolved plot threads).

What the Improved System Can Do: The system can maintain narrative fidelity over thousands of tokens by ensuring that static commitments remain inviolable while allowing dynamic elements to evolve. It prevents memory drift where a character's established trait is forgotten simply because the context window is too large.

Improvement: The system must cease treating scene generation as a freeform decoding process. Each generated scene (Scene t) must be modeled as a verifiable transition (Pre t to Post t). The system forces the LLM to output not just the text, but also a structured transition operator defining:

  • Preconditions (Pre t): What must be true in State t-1.

  • Postconditions (Post t): What must be true after generating the scene.

  • Forbidden Actions (Forbid t):: What is logically impossible given the current state.

The system then rigorously checks if State t-1 satisfies Pre t.

Improvement: The validation process must be layered:

  • Structured Validation: A deterministic check against Pre t, Post t, and Forbid t to identify explicit, hard violations (e.g., Character X cannot be in Location Y).

  • Uncertainty Monitoring: The system calculates a token-level risk score (rho i) based on the LLM's internal log probabilities for the tokens in a scene. This identifies segments that are weakly grounded or highly ambiguous, even if they don're not explicitly violating a hard rule.

Improvement: When a violation is detected (either explicitly or via high risk), the system must avoid full scene regeneration. Instead, it uses localized repair:

  • The system identifies the minimal set of conflict-bearing sentences using violation anchors.

*It applies a targeted rewriting objective (Minimize Dist(y,) subject to Violation(y) = 0).

Improvement: The system implements a strict commit policy. Only after the generated scene has been verified (i.e., Violation = 0) is the dynamic memory updated (M dt+1 = M dt M).

Sources

Related papers