ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation
Gaetano Perrone, Simon Pietro Romano
cs.CL, cs.AI
Submitted: 2026-07-31
Code: https://github.com/facebookresearch/hydra
Project page: http://skylion007.github.io/OpenWebTextCorpus
License: http://creativecommons.org/licenses/by/4.0/
The gist: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs).
Terminology
Abstract
Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it remains unclear whether performance measured on this conventional benchmark predicts detector behavior when human-authored content is rewritten by an LLM. To address this gap, we introduce Authorship-Rewriting Benchmark (ARB), built from 1,800 human source texts (600 each from XSum, WritingPrompts, and OpenWebText) and four open-weight generators (Llama-3.2-3B, Qwen2.5-7B, Mistral-7B, Gemma-2-9B). Each source item yields four matched variants: human-written (HUMAN), direct LLM generation (Free-LLM), LLM-rewritten human text (H2L), and same-generator LLM-rewritten LLM text (LLM2L). We evaluated five detectors (FastDetectGPT, Binoculars-falcon-7b, RADAR, BERT-Defense, RoBERTa-Defense) at a strict 1%-false-positive operating point (TPR@1%FPR). FastDetectGPT and Binoculars-falcon-7b detected 91.2% and 93.5% of direct LLM text, but only 30.8% and 15.1% of human text an LLM had rewritten, a drop of 60-78 percentage points. The same detectors retained 78.3% and 83.0% recall when LLM text was rewritten by the same model, a much smaller decline of 10-13 points. RADAR followed the same pattern (66.8% to 12.2%), while BERT-Defense and RoBERTa-Defense stayed below 3% recall across all regimes. These results show that detector performance measured on the conventional human-vs-LLM benchmark does not transfer to human-authored text revised by an LLM, even though the same detectors remain largely robust to LLM-only rewriting.
Sources
- Beyond Easy Wins: A Text Hardness-Aware Benchmark for LLM-generated Text Detection
- Detecting the Machine: A Comprehensive Benchmark of AI-Generated Text Detectors Across Architectures, Domains, and Adversarial Conditions
- IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text Detector
- 'Your AI Text is not Mine': Redefining and Evaluating AI-generated Text Detection under Realistic Assumptions
- The Llama 3 Herd of Models
- Mistral 7B
- Feature-Augmented Transformers for Robust AI-Text Detection Across Domains and Generators
- Gemma: Open Models Based on Gemini Research and Technology
- Benchmarking of LLM Detection: Comparing Two Competing Approaches
- Qwen2.5 Technical Report
- StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
- Can AI-Generated Text be Reliably Detected?
- A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
- The Prompt Report: A Systematic Survey of Prompt Engineering Techniques
- Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
- DUPE: Detection Undermining via Prompt Engineering for Deepfake Text
- PADBen: A Comprehensive Benchmark for Evaluating AI Text Detectors Against Paraphrase Attacks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering