A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence
Sharon Temtsin, David Kaber, Diane Proudfoot, Christoph Bartneck
University of Canterbury · Oregon State University
cs.HC, cs.AI, cs.CY
Submitted: 2026-08-08
Updated: 2026-08-11
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 66/100
The gist: The study "A RIGOROUS TURING TEST: A FOUNDATION FOR EVALUATING ARTIFICIAL GENERAL INTELLIGENCE" investigates whether state-of-the-art large language models (LLMs) can pass the Turing Test when the
Terminology
Summary
The study A RIGOROUS TURING TEST: A FOUNDATION FOR EVALUATING ARTIFICIAL GENERAL INTELLIGENCE
investigates whether state-of-the-art large language models (LLMs) can pass the Turing Test when the original instructions provided by Alan Turing are strictly followed. The researchers aimed to answer the question: Can GPT-4-Turbo, a representative of the GPT family, today’s most advanced LLMs, pass the Turing test?
and to establish guidelines for executing Turing’s instructions to support future testing.
Methodology
The researchers conducted Turing’s three-player imitation game with an LLM by following the guidelines identified by Turing and applying scientific standards wherever detailed instructions were missing.
The experiment involved two types of games: a Computer-Imitates-Human Game (CIHG)
and a man-imitates-woman game (MIWG)
used as a benchmark. Unlike previous studies that utilized time constraints, this study utilized an unconstrained duration.
The participants included 185 humans who acted as Interrogators. For the CIHG, the computer player was GPT-4-Turbo,
integrated via the OpenAI API. The model was provided with a specific prompt to simulate human thinking and typing time
and to frequently, make human-style grammatical errors and omit needless punctuation,
with instructions to deny [being a robot] by using only one sentence
and to not act like an assistant.
The setup utilized two separate rooms, one for the Contestants and one for the Interrogator,
with communication occurring through computer-mediated communication.
Results
The results indicated that GPT-4-Turbo did not pass the test.
In the CIHG, only one participant misidentified the large language model,
which indicat[es] that claims of large language models’ passing the Turing test are premature.
Statistical analysis showed that the Interrogator’s accuracy in the CIHG was significantly above the chance level,
with a mean success rate of 97%.
In comparison, the performance of the interrogators in the MIWG was 36-out-of 37, 97%.
Regarding game duration, the study found that the MIWG duration, however, was significantly longer than the CIHG duration.
The mean duration for the CIHG was 821 seconds (14 minutes),
while the mean duration for the MIWG was 1439 seconds (24 minutes).
A "generalized Wilcoxon test [50] indicated that MIWG trials lasted significantly longer than the CIHG trials (p<0.001). However, the
game duration did not appear to have a statistically significant effect on the Interrogator’s accuracy."
Discussion and Conclusion
The authors conclude that GPT-4-Turbo did not pass the test,
a finding that challenges prior claims that a digital computer passed the test.
The study suggests that previous successes reported in other studies may be due to imposed time constraint[s]
or social engineering of the prompt.
The researchers note that the popular 5-minute imitation game duration can create time pressure, which may lead human players to adopt less complex strategies.
The paper suggests that using the MIWG as a benchmark is closer to Turing’s original ideas
and recommends that future research include it. Ultimately, the authors posit that by defining a rigorous methodology for the Turing test, it is expected that the media and general public may be better able to judge the merits of future studies on 'machine intelligence'.
Improvements for AI systems
1. Non-Assistant Behavioral Framework
The system will suppress helpful assistant
heuristics—such as excessive politeness, structured list-making, and apologetic tones—and replace them with a persona-driven model. The improved AI will be able to exhibit human-like social variability, including indifference, conversational tangents, or even mild irritability, preventing the overly helpful
pattern that identifies it as a machine.
2. Cognitive Latency and Kinetic Typing Simulation
The system will implement a dynamic response delay mechanism that calculates thinking time
based on the semantic complexity of an interrogator's question. The improved AI will be able to simulate variable typing speeds and mid-sentence pauses, mimicking the fluctuating motor-skill patterns and cognitive load of a human typing in real-time.
3. Contextual Linguistic Noise Injection
The system will integrate a stochastic layer for naturalistic errors
that moves beyond simple typos. The improved AI will be able to produce context-aware grammatical slips, inconsistent punctuation, and colloquial contractions that evolve based on the duration of the conversation, simulating human fatigue or the casualness of a long-form chat.
4. High-Entropy Conversational Endurance
The system will be optimized for non-linear, meandering dialogue rather than efficient information transfer. The improved AI will be able to sustain high-engagement, low-efficiency conversations that match the 24-minute human benchmark, allowing for dead air,
circular reasoning, and topic shifts that mimic human conversational depth and duration.
Abstract
Several studies claim that large language models have passed the Turing Test and hence can "think", yet none follow Turing's original instructions precisely. Passing the test holds significance as evidence that a machine demonstrates human-like intelligence, and as a marker for artificial-general intelligence in commercial and legal domains. We conducted Turing's three-player imitation game with an LLM by following the guidelines identified by Turing and applying scientific standards wherever detailed instructions were missing. We performed a computer-imitates-human game without duration constraints and a man-imitates-woman game as a benchmark. In the computer-imitates-human game, only one participant misidentified the large language model, indicating that claims of large language models' passing the Turing test are premature. Participants required over five minutes for both tasks, with the man-imitates-woman game taking longer; shorter time limits elsewhere may explain earlier positive results. We expect the Turing Test to remain central in assessing machine intelligence for years to come.
Sources
- Human or Not? A Gamified Approach to the Turing Test
- The Turing Test Is More Relevant Than Ever
- If Turing played piano with an artificial partner
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support