Authorship identification under domain shift: a survey of stylistic measures and learned author representations
cs.CL
Submitted: 2023-09-30
Updated: 2026-09-20
License: http://creativecommons.org/licenses/by/4.0/
The gist: Authorship identification uses patterns in writing to infer who wrote a text, but those patterns also reflect topic, genre, and register.
Terminology
Abstract
Authorship identification uses patterns in writing to infer who wrote a text, but those patterns also reflect topic, genre, and register. This survey argues that topic-independence is not a property of a stylistic feature but of the feature together with its encoding, its scoring rule, and the evaluation split, so the question is which combinations preserve author differences when the domain changes. We organize the evidence in four layers, domains and evaluation conditions, linguistic measures, representations, and scoring and decision rules, and examine within-study comparisons from classical frequency measures to learned author representations and language-model-mediated writing. These comparisons show why feature inventories alone do not explain performance: normalization and encoding change results, author-supervised learning can retain content, and evaluation protocols can reverse model rankings. Successful identification, stylistic similarity, and generalization therefore require separate evidence. The literature offers a practical way to obtain it through designs that separate author identity from particular circumstances of writing: the same authors writing across genres, different authors writing within the same conversation, and parallel texts with controlled content. These designs also help evaluate content masking and language-model rewriting, whose effects depend on what identifying information remains. We use this evidence to guide the choice of stylistic representations and the comparisons needed to establish where they generalize.
Sources
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Efficient Estimation of Word Representations in Vector Space
- DT-grams: Structured Dependency Grammar Stylometry for Cross-Language Authorship Attribution
- Character-level and Multi-channel Convolutional Neural Networks for Large-scale Authorship Attribution
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering