Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking".
Jane: The paper was written by Ziwei Zhang, Juan Wen, Wanli Peng, Zhengxian Wu, Yinghan Zhou et al. from China Agricultural University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the channel, everyone. I'm Tom, and I've got Jane here with me, and we are looking at a paper that just landed on arXiv called "Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking." Jane, I have to say, the title alone got me excited, because we've talked a lot about AI stealing images, but creative writing has been this huge blind spot.
Jane: Absolutely, Tom. And I love that the title says "implicit watermarking," because that's the clever part. Normally, when we think of a watermark, we think of something you can see, like a logo stamped on a photo. But you can't stamp a watermark onto a story without ruining the story itself. So this paper is asking, how do you protect the *essence* of someone's writing without changing a single word?
Tom: Right, and that's the core problem. The authors are from China Agricultural University, which is interesting, and they're tackling the fact that anyone can feed a few chapters of a novel into an LLM and say, "write more in this style." The model picks up the voice, the rhythm, the word choices, and suddenly you've got a machine producing something that feels like Stephen King, but it's not his work.
Jane: And the paper makes a really important distinction. This isn't about figuring out who wrote a text after the fact. That's authorship attribution, which is a whole different field. This is about proactively creating a signature for a protected work, so that later, if you see a suspicious AI-generated text, you can verify whether it inherited that creative essence. It's like leaving a fingerprint on an idea, not on a sentence.
Tom: Exactly. And the way they do it is by breaking down "creative essence" into five concrete pieces. Vocabulary and word choice, syntax and grammar, rhetorical devices, tone and sentiment, and rhythm and flow. I love that they didn't just say "style" and leave it vague. They actually decomposed it into things you can measure.
Jane: That decomposition is what makes the whole thing work, Tom. Because once you have those five elements, you can ask a large language model to extract them from a piece of writing, and that becomes a condensed signature. It's like taking a whole symphony and writing down just the melody line. You lose the orchestration, but you keep what makes it recognizable.
Tom: And that's the hook for me. This could change how we think about copyright in the age of generative AI. It's not about blocking the models from reading the text. It's about giving authors a way to prove that a machine copied their *voice*, even if it didn't copy their sentences. Jane, I think we need to get into how they actually build this watermark, because the methodology is wild.
Jane: We absolutely do, because the next segment is all about the summary of the paper and the actual framework they built. Stick around, because this is where it gets technical, but I promise we'll keep it grounded.
Summary: Tom: So we're back, and we're still on "Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking." Jane, we teased the five elements, but the real meat is in the framework they call WIND, which stands for Watermarking via Implicit and Non-disruptive Disentanglement.
Jane: And that name is actually perfect, Tom. "Non-disruptive" means they don't touch the original text at all. This is what's called a zero-watermarking framework. You don't inject anything into the story. Instead, you extract a signature *from* the story and store it separately. That's a huge deal because any modification to creative writing can destroy the fragile qualities that make it good.
Tom: Right, and the way they extract that signature is the clever part. They use an LLM, like GPT-three point five or Grok, to read a piece of protected writing and produce what they call a "condensed-list." That list captures those five elements we mentioned. But here's the twist, Jane. They don't just feed the text to the LLM and hope for the best. They have this thing called an "instance delimitation mechanism."
Jane: Which is a mouthful, but the idea is simple. Before they ask the LLM to analyze a text, they first figure out which other texts are most similar to it. They use a sentence encoder to find the nearest neighbor for each sample. If the nearest neighbor is also from the protected style, they pair them up and give the LLM both texts. That gives the LLM a reference point, so it knows what to focus on.
Tom: And that matters because LLMs are sensitive to prompts. If you just say "analyze this style," you might get generic answers. But if you say "here are two texts that share a style, tell me what they have in common," the model does a much better job of isolating the creative essence from the random content. The paper actually shows that removing this delimitation step drops the F1 score by about nine points.
Jane: Nine points is significant. And once they have those condensed-lists, they map them into a watermark space. They use a learnable matrix to turn each list into a binary string, a sequence of ones and zeros. Texts that imitate the protected style get pushed toward a fixed "anchor" string, and unrelated texts get pushed toward the opposite string. That way, verification is just a matter of checking the Hamming distance.
Tom: And the results are impressive. On their main experiments, they're hitting over ninety-eight percent F1 scores, with false positive rates down around one percent. That means when they say a text is imitating the protected style, they're almost always right. And they compared against fine-tuned classifiers like BERT and RoBERTa, and WIND just blows them out of the water, especially when you only have ten or twenty examples to work with.
Jane: That few-shot capability is what makes this practical, Tom. You don't need a massive corpus of an author's work to protect them. Ten samples is enough. And that's the summary of the paper in a nutshell. But what I really want to dig into next is how they prove this thing is robust. Because anyone can build a detector that works in a lab. The question is, does it survive an attacker trying to break it?
Tom: And that's exactly the next segment, Jane. We're going to talk about the robustness studies and the attacks they threw at this system. Stay with us.
Improvements: Tom: Welcome back. We're still on "Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking," and Jane, we just said the word "robustness." So let's talk about what happens when an attacker tries to mess with the watermark.
Jane: Right, and this is where the paper gets really fun, Tom. They tested WIND against a bunch of common text attacks. Things like swapping letter cases, introducing misspellings, inserting random numbers, adding extra paragraph breaks, and even having another LLM rewrite the text entirely. The idea is that an attacker might try to disguise an imitation by making small edits that don't change the creative essence but might confuse a detector.
Tom: And the results are honestly surprising. Even under those attacks, the F1 scores stay above ninety-five percent. The rewrite attack is the toughest one, which makes sense, because rewriting changes the surface text a lot. But even then, they're at ninety-five point four eight percent F1, which is still really strong. The watermark is tied to the creative essence, not the exact wording, so superficial changes don't break it.
Jane: And that's the key improvement this paper suggests over existing methods. The old watermarking techniques, like KGW or SynthID, work by modifying the token generation process of an LLM. That means they can only protect text that the LLM itself generates with that watermark. But WIND doesn't care about the generating model at all. It works on any text, from any model, because it's extracting the style, not checking for an injected signal.
Tom: That's a massive difference. And they also did an ablation study, which is where you remove one piece of the system to see how much it matters. They found that removing the watermark loss, the thing that pushes imitations toward the anchor, drops the F1 score to seventy-three percent. That's a huge drop. And removing the instance delimitation, which we talked about earlier, drops it to about eighty-nine percent.
Jane: So every piece matters, but the watermark loss is the backbone. And they also tested different combinations of the five creative elements. Using all five is always better than using a subset. For example, on the Shakespeare dataset, using just syntax and tone gets you ninety-four point two seven percent F1, but adding vocabulary and word choice pushes it to ninety-seven point two one percent. Each element adds a little more coverage.
Tom: And they even tested cross-model generalization. They trained on texts generated by GPT-three point five and tested on texts generated by Grok, and vice versa. The performance stays high, around ninety-six to ninety-seven percent F1. That's important because in the real world, you don't know which LLM an attacker is going to use. WIND doesn't lock onto a specific model's quirks. It locks onto the creative style.
Jane: And that's the improvement that matters most, Tom. This isn't just a lab experiment. It's a system that could actually be deployed to protect authors right now. But before we wrap up, I want to hear what you think the bigger picture is. Where does this leave us in the fight between creators and generative AI?
Tom: That's the perfect question for our final segment, Jane. We're going to step back and talk about what this means for the world, and then we'll say goodbye to this paper. Don't go anywhere.
Conclusion: Tom: So we're at the end of our time with "Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking," and Jane, I want to zoom out. What does this paper actually change?
Jane: It changes the conversation, Tom. For a long time, copyright protection for text felt impossible because you can't stop a model from reading a book. But this paper shows that you don't have to stop the reading. You just have to make the copying verifiable. If an author registers their work with WIND, they get a watermark anchor. Later, if they suspect an AI-generated text is imitating them, they can test it against that anchor and get a clear yes or no.
Tom: And that's a legal tool as much as a technical one. Courts are already dealing with cases like The New York Times suing OpenAI. Having a verifiable, quantitative measure of style imitation could be the evidence that makes or breaks those cases. It moves the argument from "this feels like my writing" to "here is a ninety-eight percent match against my registered creative signature."
Jane: Exactly. And the fact that it works with just ten samples is huge. You don't need to have written a hundred novels to get protection. A short story collection, a blog, even a few essays could be enough. That democratizes the protection, which is important because it's not just famous authors who get imitated. Small creators, journalists, poets, they all have voices worth protecting.
Tom: And I love that the paper is honest about its limitations. They say the main weakness is the disentanglement performance. If a style is too close to another style, or if the creative essence is too subtle, the watermark might not separate them cleanly. But that's a direction for future work, not a reason to dismiss the approach.
Jane: Right, and future work could also expand the five elements. They mention integrating style-specific elements to broaden the expressive scope. Maybe for poetry you need different features than for technical writing. The framework is flexible enough to adapt.
Tom: So as we say goodbye to this paper, I think the takeaway is that creative writing now has a seat at the table. Images have had watermarks for years. Text was the wild west. WIND brings law and order to that frontier, without ever touching a single word of the author's work.
Jane: And that's the beautiful part, Tom. The author's text stays pristine. The watermark lives in a separate space, waiting to be called upon when needed. It's protection without interference. And with that, we're going to wrap up our discussion of "Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking." Thanks for listening, everyone.
Tom: Next up, we've got a paper on multimodal reasoning that we're excited to dig into. But for now, this is Tom and Jane, signing off. Keep reading, keep writing, and keep your voices your own.
Ziwei Zhang, Juan Wen, Wanli Peng, Zhengxian Wu, Yinghan Zhou, Yiming Xue
China Agricultural University
cs.CR, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 73/100
The gist: "We propose WIND (Watermarking via Implicit and Non-disruptive Disentanglement), a zero-watermarking framework that constructs an implicit and verifiable creative signature for copyright
Key concepts
- Implicit Watermarking
- A method of protecting text by creating a signature that is not visible or disruptive. Instead of stamping a logo, it extracts and encodes the underlying style elements of the writing so it can be verified later.
- Creative Essence
- The unique, measurable combination of an author's style, broken down into five elements: vocabulary/word choice, syntax/grammar, rhetorical devices, tone/sentiment, and rhythm/flow. This is what the watermark captures.
- WIND (Watermarking via Implicit and Non-disruptive Disentanglement)
- The framework used to create the watermark. It extracts a signature from text using an LLM and maps it into a binary string space, allowing verification of style imitation without modifying the original writing.
- Zero-watermarking framework
- A type of protection that does not inject any signal or modification into the original creative work. Instead, it analyzes and extracts a signature *from* the existing text.
Terminology
Summary
Summary
The paper introduces WIND (Watermarking via Implicit and Non-disruptive Disentanglement), a zero-watermarking framework designed to protect creative writing copyright against unauthorized imitation by large language models (LLMs). The authors state: We propose WIND (Watermarking via Implicit and Non-disruptive Disentanglement), a zero-watermarking framework that constructs an implicit and verifiable creative signature for copyright verification.
The problem addressed is that existing copyright protection techniques mainly focus on visual media, leaving the protection of creative writing largely unexplored.
The paper notes that LLMs enable unauthorized imitation of high-value creative works
through in-context learning, fine-tuning, and knowledge editing, leading to legal disputes such as the Authors Guild of America suing OpenAI and The New York Times filing a lawsuit against OpenAI.
The authors define creative writing signature through five key elements: (1) vocabulary and word choice (VWC), (2) syntactic structure and grammatical features (SSGF), (3) rhetorical devices and stylistic choices (RDCS), (4) tone and sentiment (TS), and (5) rhythm and flow (RF).
These elements collectively characterize the creative essence of writing beyond surface-level textual similarity.
The framework operates in three sequential phases: (1) Distance-driven Delimitation, (2) LLM-dominated Condensation, and (3) Watermark Mapping. In the first phase, an encoder Eα(·) computes word embeddings for positive samples (TP, texts imitating the protected creative writing) and negative samples (TN, unrelated writing). The authors use cosine similarity function sim(·)
to identify the most similar vector for each sample, and compute cross-entropy loss Lce and contrastive loss Lcon. They introduce an instance delimitation mechanism
to select optimal prior knowledge for each sample, constructing positive pair sets pp and negative sample sets neg based on whether the most similar sample is from TP or TN and whether similarity exceeds a threshold σ.
In the second phase, the authors employ a frozen-parameter LLM G(·) with two prompt templates, qp and qn, to extract condensed-lists
of five elements per sample. The input is either qp(tm, ts) for paired samples or qntm for single samples. They then impose a contrastive objective Lstyle on condensed-list embeddings to separate positive and negative samples.
In the third phase, condensed-lists are transformed into watermark vectors using a learnable watermark matrix Mγ and sigmoid function θ(·): wm = θ(Mγ · Eα(cm))
. The authors define a fixed secret anchor a ∈ 0,1 len and its complement ā = 1 − a. Positive watermarks are supervised to match anchor a, while negative watermarks are pushed toward ā, using Binary Cross-Entropy loss. The total loss is L = Lce + Lcon + λcLstyle + Lw.
For watermark validation, given a suspicious text ttest, the system identifies the most similar sample from the training dataset, extracts its condensed list, maps it to a watermark vector, and computes the probability P(wtesta) that the watermark matches the anchor. The verification score is defined as: "1, if dh(D(Ttest), D(TP)) < ϵ; 0, otherwise," where dh denotes Hamming distance and ϵ is empirically 1% of the watermark length.
Experiments were conducted using two stylistically distinct human-written datasets as protected creative writing: Shakespeare (SP) and ROCStories (ROC), with IMDB as negative creative works. Imitation texts were generated using GPT-3.5-turbo-16k, Grok-beta1, and OPT-1.3B. The encoder used was SimCSE-RoBERTa with 356.41M parameters. Training on the SP dataset with 10 samples takes approximately 26 minutes on an Apple M1 Pro chip.
Main results show that WIND outperforms fine-tuned classifiers (BERT, RoBERTa, T5) and state-of-the-art watermarking methods (KWG, Unigram, EWD, SynthID, Unbiased, MorphMark). With just 10 protected style samples, WIND achieves over 98% F1 scores while maintaining low false-positive rates.
For example, on ROC with 20 samples, WIND-3.5 achieves F1 of 97.42, TPR of 99.33, and FPR of 4.67. On SP with 20 samples, WIND-D achieves F1 of 98.47, TPR of 97.96, and FPR of 1.02.
Robustness studies against attacks including case swapping, misspellings, number insertions, adding paragraphs, and rewriting show WIND maintains strong performance. For instance, under Rewrite
attack, F1 drops only to 95.48 from 98.47.
Ablation studies demonstrate that removing any component degrades performance. Removing the delimitation mechanism (row '−qp') reduces F1 by about 9.20%. The five creative elements are shown to be complementary and non-redundant, with omitting any component leading to performance degradation. Cross-model generalization experiments show WIND maintains consistent performance under distribution shifts, e.g., training on GPT3.5 data and testing on Grok data achieves F1 of 96.75.
The authors conclude: We introduce WIND, a verifiable and implicit watermarking scheme to enable copyright verification of creative writing against unauthorized AI imitation.
They note the primary limitation is disentanglement performance
and suggest future work will optimize the feature space for creative attributes and integrate style-specific elements to broaden its expressive scope.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:
Improvement: Implement a five-dimensional disentanglement layer (VWC, SSGF, RDCS, TS, RF) that decomposes any text into structured creative attributes before processing.
What the improved system can do:
-
Extract and quantify an author's unique creative signature across five complementary dimensions rather than treating text as a monolithic block.
-
Distinguish between surface-level content similarity and deep stylistic inheritance.
-
Generate structured
condensed-lists
that capture the essence of protected works without modifying the original text.
Improvement: Add a distance-driven pre-filter that identifies optimal reference samples before any downstream task (e.g., fine-tuning, in-context learning, or watermark extraction).
Improvement: Integrate a watermark mapping phase that projects extracted creative features into a fixed-length binary anchor space without altering the source text.
Improvement: Train the disentanglement encoder with contrastive losses (L con, L style) that enforce invariance to generation model identity.
Improvement: Use the instance delimitation + LLM condensation pipeline to enable protection with minimal labeled data.
Improvement: Incorporate the five-element disentanglement to make the watermark invariant to surface-level perturbations.
Improvement: Design dual prompt templates (positive/negative) with instance-aware selection, making the system insensitive to prompt phrasing.
The improved AI system can:
-
Proactively protect creative writing from unauthorized AI imitation without modifying the original text.
-
Verify copyright with high precision (F1 >98%) and low false positives (<2%) in few-shot settings.
-
Generalize across different generation models and attack scenarios.
-
Operate in black-box environments, making it practical for real-world legal and publishing workflows.
-
Scale to low-resource scenarios, requiring only 10–20 samples for robust protection.
Sources
- Trends in Integration of Knowledge and Large Language Models: A Survey and Taxonomy of Methods, Benchmarks, and Applications
- DeepSeek-V3 Technical Report
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Watermarking Text Data on Large Language Models for Dataset Copyright
- GPT-4 Technical Report
- A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
- OPT: Open Pre-trained Transformer Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs