Who Wrote the Book? Detecting and Attributing LLM Ghostwriters

summary

Video file (mp4)

The gist

Out-of-distribution generalization remains insufficiently supported by existing authorship attribution datasets, which often focus on short texts and fail to test for unseen authors or domains.

In short

Existing authorship attribution methods fail when tested against unseen authors or domains in long texts. This work introduces GHOSTWRITEBENCH, a new dataset testing generalization across domain and author variations for LLM writing. It proposes TRACE, a lightweight fingerprinting method that captures token-level transition patterns to establish a strong baseline for attributing text generated by large language models.

Key concepts

GHOSTWRITEBENCH Dataset
A new dataset specifically designed to test whether an LLM can attribute authorship accurately on long texts. It tests generalization across two dimensions: unseen writing styles (domain) and texts from different LLMs (author). It includes long books generated by ten frontier LLMs.
OOD-Domain Generalization
Testing if a model can correctly identify the author when the text belongs to a genre or style it has never seen before. This is tested by splitting book genres into known and unknown sets for each specific author.
TRACE Fingerprinting Method
A lightweight technique that creates a unique signature for an LLM's writing style based on how words transition from one to the next. It uses token-level patterns, such as word rank and entropy, captured by an evaluator language model, to create a compressed two-dimensional signature.
OOD-Author Setting
A challenging test where the attribution method must identify an author based on text generated by an LLM it has never been trained on. Existing methods struggle significantly here, often losing up to 90% performance compared to TRACE.

Terminology used across episodes

This episode discusses

The paper

Who Wrote the Book? Detecting and Attributing LLM Ghostwriters · Read on arXiv

School of Computing and Information Systems, The University of Melbourne · School of Computing, FSE, Macquarie University

In this paper, we introduce GhostWriteBench, a dataset for LLM authorship attribution. It comprises long-form texts (50K+ words per book) generated by frontier LLMs, and is designed to test generalisation across multiple out-of-distribution (OOD) dimensions, including domain and unseen LLM author. We also propose TRACE -- a novel fingerprinting method that is interpretable and lightweight -- that works for both open- and closed-source models. TRACE creates the fingerprint by capturing token-level transition patterns (e.g., word rank) estimated by another lightweight language model. Experiments on GhostWriteBench demonstrate that TRACE achieves state-of-the-art performance, remains robust in OOD settings, and works well in limited training data scenarios.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Who Wrote the Book? Detecting and Attributing LLM Ghostwriters".

Jane: Out-of-distribution generalization remains insufficiently supported by existing authorship attribution datasets, which often focus on short texts and fail to test for unseen authors or domains.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Moving on, the paper 'Who Wrote the Book? Detecting and Attributing LLM Ghostwriters' summarizes their main idea, which is proposing a new fingerprinting method called TRACE to solve this problem.

Jane: They explain that TRACE works by capturing token-level transition patterns, like how words follow each other or how complex the language is at specific points in the text.

Lu: TRACE uses a lightweight language model as an evaluator to create a compressed two-dimensional signature based on word rank and entropy, which they call a fingerprint.

Meng: A token-level approach sounds very detailed; how does that actually translate into a usable signature for classification? I mean, can we really process all those transitions efficiently?

Lalam: TRACE is designed to be lightweight and interpretable because it focuses on these internal transition patterns rather than just relying on heavy, opaque neural network layers.

Tom: The authors show that TRACE establishes itself as a strong baseline because it's robust in those out-of-distribution settings where other methods fail so badly.

Jane: They also present two variants of this method: one based on word rank and another based on token entropy, both aiming to create a compressed signature for comparison.

Lu: This paper is important because it moves the focus away from traditional attribution labels toward capturing these inherent stylistic characteristics unique to the LLMs.

Meng: So, the core idea is that we can build a fingerprint based on how an AI chooses its next word, regardless of what specific book it's writing.

Lalam: That’s right; they argue that this method captures LLM-specific stylistic traits rather than just content features like genre.

The paper's summary: Tom: Now we get into the suggested improvements they propose, which really shows how this research can be put to use in making attribution systems much tougher.

Jane: They suggest replacing heavier transformer-based classification backbones with the TRACE fingerprinting mechanism for a more robust and generalizable system.

Lu: The paper argues that using TRACE allows for better performance across out-of-distribution settings, especially when dealing with unseen authors or new domains.

Meng: If we adopt this idea, it means we can fine-tune existing classifiers like BERT-AA on GHOSTWRITEBENCH and see a much bigger improvement in resilience against novel content.

Lalam: They also propose using the Entropy-based variant of TRACE, which creates a continuous density map through kernel density estimation, which they claim captures model-specific stylistic patterns very effectively.

Tom: That continuous map idea is interesting; it suggests we aren't just looking for hard boundaries but understanding the overall shape of how an AI writes.

Jane: They also suggest that by using GHOSTWRITEBENCH to stress-test models across different splits, we force them to develop these intrinsic fingerprinting abilities.

Lu: This work pushes the idea that attribution systems should be designed around capturing these stylistic characteristics rather than just training them on specific author labels.

Meng: From a practical standpoint, this implies that if we can implement TRACE, it makes our AI detection tools way more reliable when they encounter text from a source they haven't seen before.

Lalam: This is crucial because it builds resilience into the system itself by focusing on how the AI generates text at the token level rather than relying solely on supervision.

The paper's improvements: Tom: So, to wrap up, the conclusion of 'Who Wrote the Book? Detecting and Attributing LLM Ghostwriters' boils down to confirming that TRACE is a strong baseline for LLM attribution tasks.

Jane: They emphasize that this method generalizes because it relies on token-level transition patterns computed by an independent evaluator language model, instead of direct supervision from attribution labels.

Lu: This suggests the power comes from capturing those internal stylistic characteristics of the LLM's output, making it less dependent on specific training data.

Meng: If we consider the practical impact, this means our future AI systems could be much more dependable when trying to verify content integrity in a vast and evolving digital landscape.

Lalam: This paper sets a new benchmark for text forensics by showing that capturing these granular transition patterns can lead to better detection methods overall.

Tom: It’s a solid piece of research that gives us a clear path forward for creating more resilient AI systems capable of handling novel writing styles.

Jane: It really shows how focusing on the mechanics of token generation, like word rank and entropy, provides a stable foundation for authorship classification.

Lu: I think the most exciting part is how this methodology could evolve into a way to model not just who wrote the text, but also how different LLMs develop their writing styles over time.

Meng: That's a big thought; keeping track of those evolving styles would be essential for any large-scale content moderation or verification system.

Lalam: It’s exciting to think about how this approach can help us understand the cultural impact of AI-generated content by identifying the underlying patterns.

Conclusion: Tom: So we’ve been diving deep into "Who Wrote the Book? Detecting and Attributing LLM Ghostwriters," and what we really learned is that TRACE offers a solid, robust baseline for tackling authorship attribution in these long-form text scenarios.

Jane: That's right, Tom; they showed how capturing token-level transitions through word rank and entropy gives us a way to fingerprint an AI's writing style without needing heavy supervision.

Lu: I think the real magic here is moving away from just looking at content and focusing on these underlying generative patterns; it opens up so many creative avenues for understanding LLM behavior.

Meng: From an engineering standpoint, this means we can potentially build more reliable systems that don't rely on specific author names but instead identify the AI model itself based on its unique linguistic DNA.

Lalam: I see the immense cultural potential here; if we can reliably detect when text is ghostwritten, it helps us establish clearer guidelines for digital creativity and authenticity in our media landscape.

Tom: Exactly, Lalam; it’s about building a foundation for trust in AI-generated content. The results on GHOSTWRITEBENCH were quite striking, showing how much existing methods struggle under those out-of-distribution conditions.

Jane: It really highlights why this work is important; traditional attribution methods just don't have the flexibility to handle completely new authors or completely different genres of writing effectively.

Lu: The paper’s methodology with both rank-based and entropy-based fingerprints gives us two distinct lenses through which to analyze these transition patterns, which could lead to even more nuanced detection strategies down the road.

Meng: I wonder how much computational overhead TRACE adds compared to a full transformer model, but if it stays lightweight like they claim, it could be very practical for real-time text scanning applications.

Lalam: The ability to generalize across unseen authors and domains is huge; it means the system isn't just memorizing patterns from the training set but actually learning what makes an AI *write* in a fundamental way.

Tom: It’s certainly a big step forward in moving past simple pattern matching toward understanding the core stylistic choices made by these language models.

Jane: And it sets a very clear direction for future research, showing us exactly where the current limitations are and where the next generation of attribution tools needs to focus their efforts.

Lu: This paper provides a fantastic starting point for exploring how we can build systems that truly understand the generative process rather than just classifying static text features.

Meng: I'm looking forward to seeing how this TRACE mechanism integrates into larger, more complex AI pipelines in the coming months.

Lalam: The implications for culture are vast because this gives us a tool that could help distinguish genuine human expression from sophisticated machine mimicry at an unprecedented level of detail.

Tom: Absolutely, it’s a compelling study on how we can build better tools to navigate the complexities of AI-generated text. We'll be back after the break with more deep dives into new arXiv papers!

More episodes

← Home