Who Wrote the Book? Detecting and Attributing LLM Ghostwriters

arXiv:2603.28054 · cs.CL · Submitted 2026-03-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Who Wrote the Book? Detecting and Attributing LLM Ghostwriters".

Jane: Out-of-distribution generalization remains insufficiently supported by existing authorship attribution datasets, which often focus on short texts and fail to test for unseen authors or domains.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Moving on, the paper 'Who Wrote the Book? Detecting and Attributing LLM Ghostwriters' summarizes their main idea, which is proposing a new fingerprinting method called TRACE to solve this problem.

Jane: They explain that TRACE works by capturing token-level transition patterns, like how words follow each other or how complex the language is at specific points in the text.

Lu: TRACE uses a lightweight language model as an evaluator to create a compressed two-dimensional signature based on word rank and entropy, which they call a fingerprint.

Meng: A token-level approach sounds very detailed; how does that actually translate into a usable signature for classification? I mean, can we really process all those transitions efficiently?

Lalam: TRACE is designed to be lightweight and interpretable because it focuses on these internal transition patterns rather than just relying on heavy, opaque neural network layers.

Tom: The authors show that TRACE establishes itself as a strong baseline because it's robust in those out-of-distribution settings where other methods fail so badly.

Jane: They also present two variants of this method: one based on word rank and another based on token entropy, both aiming to create a compressed signature for comparison.

Lu: This paper is important because it moves the focus away from traditional attribution labels toward capturing these inherent stylistic characteristics unique to the LLMs.

Meng: So, the core idea is that we can build a fingerprint based on how an AI chooses its next word, regardless of what specific book it's writing.

Lalam: That’s right; they argue that this method captures LLM-specific stylistic traits rather than just content features like genre.

The paper's summary: Tom: Now we get into the suggested improvements they propose, which really shows how this research can be put to use in making attribution systems much tougher.

Jane: They suggest replacing heavier transformer-based classification backbones with the TRACE fingerprinting mechanism for a more robust and generalizable system.

Lu: The paper argues that using TRACE allows for better performance across out-of-distribution settings, especially when dealing with unseen authors or new domains.

Meng: If we adopt this idea, it means we can fine-tune existing classifiers like BERT-AA on GHOSTWRITEBENCH and see a much bigger improvement in resilience against novel content.

Lalam: They also propose using the Entropy-based variant of TRACE, which creates a continuous density map through kernel density estimation, which they claim captures model-specific stylistic patterns very effectively.

Tom: That continuous map idea is interesting; it suggests we aren't just looking for hard boundaries but understanding the overall shape of how an AI writes.

Jane: They also suggest that by using GHOSTWRITEBENCH to stress-test models across different splits, we force them to develop these intrinsic fingerprinting abilities.

Lu: This work pushes the idea that attribution systems should be designed around capturing these stylistic characteristics rather than just training them on specific author labels.

Meng: From a practical standpoint, this implies that if we can implement TRACE, it makes our AI detection tools way more reliable when they encounter text from a source they haven't seen before.

Lalam: This is crucial because it builds resilience into the system itself by focusing on how the AI generates text at the token level rather than relying solely on supervision.

The paper's improvements: Tom: So, to wrap up, the conclusion of 'Who Wrote the Book? Detecting and Attributing LLM Ghostwriters' boils down to confirming that TRACE is a strong baseline for LLM attribution tasks.

Jane: They emphasize that this method generalizes because it relies on token-level transition patterns computed by an independent evaluator language model, instead of direct supervision from attribution labels.

Lu: This suggests the power comes from capturing those internal stylistic characteristics of the LLM's output, making it less dependent on specific training data.

Meng: If we consider the practical impact, this means our future AI systems could be much more dependable when trying to verify content integrity in a vast and evolving digital landscape.

Lalam: This paper sets a new benchmark for text forensics by showing that capturing these granular transition patterns can lead to better detection methods overall.

Tom: It’s a solid piece of research that gives us a clear path forward for creating more resilient AI systems capable of handling novel writing styles.

Jane: It really shows how focusing on the mechanics of token generation, like word rank and entropy, provides a stable foundation for authorship classification.

Lu: I think the most exciting part is how this methodology could evolve into a way to model not just who wrote the text, but also how different LLMs develop their writing styles over time.

Meng: That's a big thought; keeping track of those evolving styles would be essential for any large-scale content moderation or verification system.

Lalam: It’s exciting to think about how this approach can help us understand the cultural impact of AI-generated content by identifying the underlying patterns.

Conclusion: Tom: So we’ve been diving deep into "Who Wrote the Book? Detecting and Attributing LLM Ghostwriters," and what we really learned is that TRACE offers a solid, robust baseline for tackling authorship attribution in these long-form text scenarios.

Jane: That's right, Tom; they showed how capturing token-level transitions through word rank and entropy gives us a way to fingerprint an AI's writing style without needing heavy supervision.

Lu: I think the real magic here is moving away from just looking at content and focusing on these underlying generative patterns; it opens up so many creative avenues for understanding LLM behavior.

Meng: From an engineering standpoint, this means we can potentially build more reliable systems that don't rely on specific author names but instead identify the AI model itself based on its unique linguistic DNA.

Lalam: I see the immense cultural potential here; if we can reliably detect when text is ghostwritten, it helps us establish clearer guidelines for digital creativity and authenticity in our media landscape.

Tom: Exactly, Lalam; it’s about building a foundation for trust in AI-generated content. The results on GHOSTWRITEBENCH were quite striking, showing how much existing methods struggle under those out-of-distribution conditions.

Jane: It really highlights why this work is important; traditional attribution methods just don't have the flexibility to handle completely new authors or completely different genres of writing effectively.

Lu: The paper’s methodology with both rank-based and entropy-based fingerprints gives us two distinct lenses through which to analyze these transition patterns, which could lead to even more nuanced detection strategies down the road.

Meng: I wonder how much computational overhead TRACE adds compared to a full transformer model, but if it stays lightweight like they claim, it could be very practical for real-time text scanning applications.

Lalam: The ability to generalize across unseen authors and domains is huge; it means the system isn't just memorizing patterns from the training set but actually learning what makes an AI *write* in a fundamental way.

Tom: It’s certainly a big step forward in moving past simple pattern matching toward understanding the core stylistic choices made by these language models.

Jane: And it sets a very clear direction for future research, showing us exactly where the current limitations are and where the next generation of attribution tools needs to focus their efforts.

Lu: This paper provides a fantastic starting point for exploring how we can build systems that truly understand the generative process rather than just classifying static text features.

Meng: I'm looking forward to seeing how this TRACE mechanism integrates into larger, more complex AI pipelines in the coming months.

Lalam: The implications for culture are vast because this gives us a tool that could help distinguish genuine human expression from sophisticated machine mimicry at an unprecedented level of detail.

Tom: Absolutely, it’s a compelling study on how we can build better tools to navigate the complexities of AI-generated text. We'll be back after the break with more deep dives into new arXiv papers!

School of Computing and Information Systems, The University of Melbourne · School of Computing, FSE, Macquarie University

cs.CL

Submitted: 2026-03-30

Updated: 2026-10-01

Code: https://github.com/anudeex/TRACE

Importance score: 86/100

The gist: Out-of-distribution generalization remains insufficiently supported by existing authorship attribution datasets, which often focus on short texts and fail to test for unseen authors or domains.

Key concepts

GHOSTWRITEBENCH Dataset
A new dataset specifically designed to test whether an LLM can attribute authorship accurately on long texts. It tests generalization across two dimensions: unseen writing styles (domain) and texts from different LLMs (author). It includes long books generated by ten frontier LLMs.
OOD-Domain Generalization
Testing if a model can correctly identify the author when the text belongs to a genre or style it has never seen before. This is tested by splitting book genres into known and unknown sets for each specific author.
TRACE Fingerprinting Method
A lightweight technique that creates a unique signature for an LLM's writing style based on how words transition from one to the next. It uses token-level patterns, such as word rank and entropy, captured by an evaluator language model, to create a compressed two-dimensional signature.
OOD-Author Setting
A challenging test where the attribution method must identify an author based on text generated by an LLM it has never been trained on. Existing methods struggle significantly here, often losing up to 90% performance compared to TRACE.

Terminology

Summary

Out-of-distribution generalization remains insufficiently supported by existing authorship attribution datasets, which often focus on short texts and fail to test for unseen authors or domains. This paper introduces GHOSTWRITEBENCH, a new dataset designed to test generalisation across multiple out-of-distribution dimensions (OOD), and proposes TRACE, a novel lightweight fingerprinting method that captures token-level transition patterns to establish a strong baseline for LLM attribution.

GHOSTWRITEBENCH Dataset Construction

The paper introduces GHOSTWRITEBENCH, a dataset specifically designed for LLM authorship attribution (AA) that supports long-form book detection and explicitly tests generalisation across two OOD dimensions: OOD-Domain (unseen genres) and OOD-Author (unseen LLMs). The dataset comprises long-form texts (50K+ words per book) generated by ten frontier LLMs. The generation pipeline mimics the human process involving outlining, summarisation, and long-text expansion, using real, human-authored texts from Project Gutenberg as few-shot reference examples to guide the LLM generation.

OOD Dimensions and Scenarios

GHOSTWRITEBENCH supports two primary OOD dimensions:

  1. OOD-Domain: Testing generalisation to texts written in unseen genres. This is achieved by partitioning book genres into ID and OOD sets for each author, creating per-author domain splits.

  2. OOD-Author: Testing generalisation to texts generated by unseen LLMs. This is constructed by holding out one LLM at a time, where the training partition contains books from nine LLMs and the test partition contains books from the held-out LLM, reporting performance across ten such splits.

Furthermore, GHOSTWRITEBENCH provides two training scenarios within each evaluation setting: high-resource and low-resource (where low-resource involves a small number of training books (3–5) text–label pairs per LLM). The paper notes that existing attribution methods struggle under the OOD-Author setting, with strong methods suffering drops of up to 90% in low-resource and 70% in high-resource training scenarios.

TRACE Fingerprinting Method

The paper proposes TRACE (Transition-based Representation for Authorship Classification and Evaluation), a novel fingerprinting method that is lightweight and interpretable. TRACE first constructs reference fingerprints for each LLM using a (small) set of training texts generated by an LLM. At test time, it creates the test fingerprint and determines attribution by assessing its similarity to the reference fingerprints.

TRACE examines token-level transition patterns (i.e., word rank and entropy) using a lightweight language model as an evaluator. The method captures these transitions to create a compressed two-dimensional signature. Two variants are proposed:

  1. Rank-based: This involves computing the rank of each token under an evaluator language model, creating a raw fingerprint matrix that is then compressed by binning ranks into equal probability clusters.

  2. Entropy-Based: This calculates the entropy of a token, and then uses kernel density estimation (KDE) to obtain a continuous density map, resulting in an entropy-based fingerprint.

Experimental Results and Performance

Experiments on GHOSTWRITEBENCH show that TRACE establishes itself as a strong baseline on GHOSTWRITEBENCH. The results demonstrate the robustness of TRACE in OOD settings. For instance, Table 3 shows that under the OOD-Author setting, existing methods suffer significant degradation, while TRACE achieves high performance. Specifically, in the OOD-Author scenario for low-resource training settings, TRACE's Entropy-JS score is reported as 1.00 ± 0.00, achieving near-perfect performance against unseen human authors in the OOD-Author variant (Table 4).

Conclusion and Contribution

The primary contributions are:

  1. Introducing GHOSTWRITEBENCH, a new dataset for LLM authorship attribution that supports long-form book detection and explicitly tests generalisation across two OOD dimensions (OOD-Domain and OOD-Author).

  2. Proposing TRACE, a lightweight and interpretable fingerprinting technique for LLM attribution that captures token-level transition patterns.

The paper concludes by showing that TRACE generalises because it relies on token-level transition patterns computed by an independent evaluator language model, rather than direct supervision on attribution labels, suggesting it captures LLM-specific stylistic characteristics rather than content. The work establishes a new benchmark for future research in LLM text forensics.

Limitations

The limitations acknowledged include:

  1. GHOSTWRITEBENCH is limited to English texts only, with plans to extend it crosslingually.

  2. Adversarial robustness (e.g.

Improvements for AI systems

Here are specific improvements to existing AI systems based on the findings of this research, along with what those improved systems could achieve:


) 1. Improved Authorship Attribution (AA) Systems for Long-Form Content:

A novel, lightweight fingerprinting method called TRACE (Transition-based Representation for Authorship Classification and Evaluation) can be integrated into existing AA pipelines.

  • Specific Improvement: Replace or augment current heavy transformer-based classification backbones with the TRACE fingerprinting mechanism. This involves using a small, lightweight language model (like GPT-2 used in the paper) to compute token-level transition patterns (word rank and entropy) as a compressed 2D signature.

  • What it can do: It enables robust, generalizable authorship attribution for long documents (books, >10K words). Crucially, it maintains high performance across Out-of-Distribution (OOD) settings—specifically unseen authors and unseen domains—where existing methods degrade significantly. This allows the system to reliably detect LLM ghostwritten content even when the generating model is new or the writing style is novel.

    1. Enhanced Generalization in Multi-Class Classification:

Existing methods often fail when applied to new authors or domains because they rely on supervised training with known authors/domains.

  • Specific Improvement: Implement GHOSTWRITEBENCH as a benchmark for stress-testing attribution models, and fine-tune existing classifiers (e.g., BERT-AA, TOPFORMER) using this dataset across multiple OOD splits (OOD-Domain and OOD-Author).

  • What it can do: It forces the AI to develop fingerprinting capabilities based on intrinsic linguistic style rather than surface features or specific author labels. The resulting system will be far more resilient in real-world scenarios where it encounters content from sources it hasn't seen before, drastically reducing performance drops (up to 90% improvement over existing baselines in low-resource settings).

    1. Development of Domain-Agnostic Stylometric Detectors:

The TRACE method captures token transitions rather than static distributions, leading to more discriminative signatures.

  • Specific Improvement: Utilize the Entropy-based variant of TRACE, which produces a continuous density map (via Kernel Density Estimation) and is shown to capture model-specific stylistic patterns rather than genre features (as seen in Figure 2).

  • What it can do: The AI system will be able to distinguish between different LLMs based purely on their underlying generation DNA or transition preferences, irrespective of the specific content genre (e.g., Fiction vs. Science), making the detection mechanism highly versatile and less prone to genre-specific overfitting.

    1. Robust Model Selection and Evaluation:

The research demonstrates that TRACE is robust across different evaluator language models (GPT-2, OLMo, Gemma).

  • Specific Improvement: Design a modular TRACE architecture where the evaluator language model component is easily swappable without retraining the core fingerprinting mechanism.

  • What it can do: This allows for rapid adaptation of the detection system to newer or proprietary frontier models simply by plugging in a compatible evaluator model, ensuring the attribution system remains state-of-the-art as LLMs evolve.

    1. Improved Text Quality and Diversity Assessment (for Content Generation):

The paper establishes that GHOSTWRITEBENCH texts exhibit high scores across human evaluation metrics (Relevance, Coherence, Fluency).

  • Specific Improvement: Integrate the quality metrics used in GHOSTWRITEBENCH (PPL, S-B, S-R, NGD) into a feedback loop for LLM training or fine-tuning.

  • What it can do: This creates a self-correction mechanism where LLMs are rewarded not just for generating coherent text, but for generating text that exhibits high linguistic diversity and low self-repetition—qualities proven to align with high human quality scores.

    1. Advanced Long-Form Content Structuring (for LLM Generation):

The work introduces sophisticated prompt engineering techniques (Outline generation, incremental summarization) for sequential, long-form text generation.

  • Specific Improvement: Adopt the hierarchical book generation pipeline (outline -> segment -> summary update) as a standard operational procedure for any LLM tasked with writing books or complex narratives.

  • What it can do: This allows AI to generate high-quality, multi-part long documents sequentially, ensuring narrative coherence and consistency over vast word counts, overcoming the traditional context limits of standard transformer models.

Sources

Related papers