Towards Detecting AI-Assisted Responses in Online Surveys
cs.CL, cs.CY
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: Accepted to EMNLP 2026 (Main Conference)
Code: https://github.com/mike-qz-wang/ASURRE
License: http://creativecommons.org/licenses/by/4.0/
The gist: The use of LLMs to complete online surveys impacts the validity of survey-based research, but detecting such usage remains underexplored.
Terminology
Abstract
The use of LLMs to complete online surveys impacts the validity of survey-based research, but detecting such usage remains underexplored. We introduce an initial benchmark dataset, namely ASURRE, for AI-assisted survey participation to capture usage strategies ranging from full generation and revision to persona-grounded agentic completion. Controlled by these strategies, LLM-assisted survey responses are generated using multiple LLMs on three real-world surveys in different disciplines, paired with genuine human responses. Our evaluation of existing machine-generated text (MGT) detectors shows that naive AI usage is readily detectable, whereas persona-grounded agents that mimic entire respondents push detector performance toward chance. We further show that agentic completion cannot fully replicate respondent-level behaviour and leaves distinctive behavioural traces. While individual cues can be circumvented by targeted prompting, a simple few-shot, training-free aggregator over these cues improves mean AUROC by +0.14 over the best existing detector across agentic settings. Our project is available at https://github.com/mike-qz-wang/ASURRE.
Sources
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
- gpt-oss-120b & gpt-oss-20b Model Card
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
- RADAR: Robust AI-Text Detection via Adversarial Learning
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Release Strategies and the Social Impacts of Language Models
- Artificial Artificial Artificial Intelligence: Crowd Workers Widely Use Large Language Models for Text Production Tasks
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering