Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue
cs.CL, cs.CY
Submitted: 2026-06-20
Updated: 2026-09-10
License: http://creativecommons.org/licenses/by/4.0/
The gist: As AI systems integrate into online spaces, differentiating them from humans in conversations is increasingly important.
Terminology
Abstract
As AI systems integrate into online spaces, differentiating them from humans in conversations is increasingly important. We present Inverse Turing Bench, a benchmark that evaluates LLMs and other models on their ability to differentiate humans and AI in multi-turn text. The benchmark provides a collection of paired dialogue transcripts, wherein one dialogue is between two humans and the other is between a human and an AI. The task is to correctly identify which dialogue is human-only vs. human-AI. We evaluated a preliminary set of models against this benchmark, and found that GPTZero, Claude Opus-4.6, and GPT-5.5 achieve the highest accuracy: 89.41%, 77.92%, and 75.94% respectively. Our results suggest that statistical approaches to detection have semantic blind spots, but semantic approaches are susceptible to persona-prompting. Our work speaks to the Inverse Turing Test and motivates human-AI differentiation as a critical capability for AI systems. Our live benchmark can be found at https://huggingface.co/spaces/roc-hci/Inverse-Turing-Bench-Leaderboard.
Sources
- Disclosure By Design: Identity Transparency as a Behavioural Property of Conversational AI Models
- TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
- "Humans welcome to observe": A First Look at the Agent Social Network Moltbook
- Decoding the AI Pen: Techniques and Challenges in Detecting AI-Generated Text
- GPTZero: Robust Detection of LLM-Generated Texts
- The Llama 3 Herd of Models
- International AI Safety Report 2026
- OPT: Open Pre-trained Transformer Language Models
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering