Benchmark of stylistic variation in LLM-generated texts
cs.CL, cs.AI
Submitted: 2025-09-12
Updated: 2025-09-18
Comments: Data and scripts: https://osf.io/hs7xt/. Interactive charts: https://www.korpus.cz/stylisticbenchmark/
Journal ref: International Journal of Corpus Linguistics 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: This study investigates the register variation in texts written by humans and comparable texts produced by large language models (LLMs).
Terminology
Abstract
This study investigates the register variation in texts written by humans and comparable texts produced by large language models (LLMs). Biber's multidimensional analysis (MDA) is applied to a sample of human-written texts and AI-created texts generated to be their counterparts to find the dimensions of variation in which LLMs differ most significantly and most systematically from humans. As textual material, a new LLM-generated corpus AI-Brown is used, which is comparable to BE-21 (a Brown family corpus representing contemporary British English). Since all languages except English are underrepresented in the training data of frontier LLMs, similar analysis is replicated on Czech using AI-Koditex corpus and Czech multidimensional model. Examined were 16 frontier models in various settings and prompts, with emphasis placed on the difference between base models and instruction-tuned models. Based on this, a benchmark is created through which models can be compared with each other and ranked in interpretable dimensions.
Sources
- Language Models are Few-Shot Learners
- The Ghost in the Machine has an American accent: value conflict in GPT-3
- Stylometric Detection of AI-Generated Text in Twitter Timelines
- Humans can learn to detect AI-generated texts, or at least learn when they can't
- Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
- Do Multilingual LLMs Think In English?
- Beyond English-Centric LLMs: What Language Do Multilingual Language Models Think in?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering