Interpretable Predictability-Based AI Text Detection: A Replication Study
cs.CL, cs.AI, cs.LG
Submitted: 2026-03-16
Updated: 2026-08-31
Comments: Findings of EMNLP 2026
Code: https://github.com/piotrmp/autexthttps:
License: http://creativecommons.org/licenses/by/4.0/
The gist: This paper replicates and extends the system used in the AuTexTification shared task for authorship attribution of machine-generated texts.
Terminology
Abstract
This paper replicates and extends the system used in the AuTexTification shared task for authorship attribution of machine-generated texts. Exact replication was not possible because of differences in data splits, model availability, and implementation details, which we document as a case study in reproducibility. We tested newer multilingual language models (mDeBERTa-v3-base, Qwen, mGPT) and added 26 document-level stylometric features, using ablation, permutation importance, and SHAP analysis to assess feature influence. A single shared configuration was applied to both English and Spanish across Subtask 1 and Subtask 2. Averaged over three random seeds, the shared multilingual configuration performs comparably to or better than the language-specific baseline, with the clearest gains on model attribution (Subtask 2). The additional stylometric features yield small improvements, led by lexical diversity, but their contribution falls within seed variance once predictability-based probabilities are included, which remain the dominant signal. The study also shows that clear documentation is important for reliable replication and fair comparison of systems.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
- Authorship Attribution in Multilingual Machine-Generated Texts
- Unsupervised Cross-lingual Representation Learning at Scale
- GLTR: Statistical Detection and Visualization of Generated Text
- The Llama 3 Herd of Models
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- GPT-4 Technical Report
- How multilingual is Multilingual BERT?
- Adversarial Attacks on AI-Generated Text Detection Models: A Token Probability-Based Approach Using Embeddings
- Robust AI-Generated Text Detection by Restricted Embeddings
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Classification of Human- and AI-Generated Texts for English, French, German, and Spanish
- From Understanding to Utilization: A Survey on Explainability for Large Language Models
- mdok of KInIT: Robustly Fine-tuned LLM for Binary and Multiclass AI-Generated Text Detection
- Few-Shot Detection of Machine-Generated Text using Style Representations
- DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated Text
- TRACE: TRansformer-based Attribution using Contrastive Embeddings in LLMs
- M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering