A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence

arXiv:2501.17629 · cs.HC, cs.AI, cs.CY · Submitted 2026-08-08 · Read on arXiv

Sharon Temtsin, David Kaber, Diane Proudfoot, Christoph Bartneck

University of Canterbury · Oregon State University

cs.HC, cs.AI, cs.CY

Submitted: 2026-08-08

Updated: 2026-08-11

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 66/100

The gist: The study "A RIGOROUS TURING TEST: A FOUNDATION FOR EVALUATING ARTIFICIAL GENERAL INTELLIGENCE" investigates whether state-of-the-art large language models (LLMs) can pass the Turing Test when the

Terminology

Summary

The study A RIGOROUS TURING TEST: A FOUNDATION FOR EVALUATING ARTIFICIAL GENERAL INTELLIGENCE investigates whether state-of-the-art large language models (LLMs) can pass the Turing Test when the original instructions provided by Alan Turing are strictly followed. The researchers aimed to answer the question: Can GPT-4-Turbo, a representative of the GPT family, today’s most advanced LLMs, pass the Turing test? and to establish guidelines for executing Turing’s instructions to support future testing.

Methodology

The researchers conducted Turing’s three-player imitation game with an LLM by following the guidelines identified by Turing and applying scientific standards wherever detailed instructions were missing. The experiment involved two types of games: a Computer-Imitates-Human Game (CIHG) and a man-imitates-woman game (MIWG) used as a benchmark. Unlike previous studies that utilized time constraints, this study utilized an unconstrained duration.

The participants included 185 humans who acted as Interrogators. For the CIHG, the computer player was GPT-4-Turbo, integrated via the OpenAI API. The model was provided with a specific prompt to simulate human thinking and typing time and to frequently, make human-style grammatical errors and omit needless punctuation, with instructions to deny [being a robot] by using only one sentence and to not act like an assistant. The setup utilized two separate rooms, one for the Contestants and one for the Interrogator, with communication occurring through computer-mediated communication.

Results

The results indicated that GPT-4-Turbo did not pass the test. In the CIHG, only one participant misidentified the large language model, which indicat[es] that claims of large language models’ passing the Turing test are premature. Statistical analysis showed that the Interrogator’s accuracy in the CIHG was significantly above the chance level, with a mean success rate of 97%. In comparison, the performance of the interrogators in the MIWG was 36-out-of 37, 97%.

Regarding game duration, the study found that the MIWG duration, however, was significantly longer than the CIHG duration. The mean duration for the CIHG was 821 seconds (14 minutes), while the mean duration for the MIWG was 1439 seconds (24 minutes). A "generalized Wilcoxon test [50] indicated that MIWG trials lasted significantly longer than the CIHG trials (p<0.001). However, the game duration did not appear to have a statistically significant effect on the Interrogator’s accuracy."

Discussion and Conclusion

The authors conclude that GPT-4-Turbo did not pass the test, a finding that challenges prior claims that a digital computer passed the test. The study suggests that previous successes reported in other studies may be due to imposed time constraint[s] or social engineering of the prompt. The researchers note that the popular 5-minute imitation game duration can create time pressure, which may lead human players to adopt less complex strategies.

The paper suggests that using the MIWG as a benchmark is closer to Turing’s original ideas and recommends that future research include it. Ultimately, the authors posit that by defining a rigorous methodology for the Turing test, it is expected that the media and general public may be better able to judge the merits of future studies on 'machine intelligence'.

Improvements for AI systems

1. Non-Assistant Behavioral Framework

The system will suppress helpful assistant heuristics—such as excessive politeness, structured list-making, and apologetic tones—and replace them with a persona-driven model. The improved AI will be able to exhibit human-like social variability, including indifference, conversational tangents, or even mild irritability, preventing the overly helpful pattern that identifies it as a machine.

2. Cognitive Latency and Kinetic Typing Simulation

The system will implement a dynamic response delay mechanism that calculates thinking time based on the semantic complexity of an interrogator's question. The improved AI will be able to simulate variable typing speeds and mid-sentence pauses, mimicking the fluctuating motor-skill patterns and cognitive load of a human typing in real-time.

3. Contextual Linguistic Noise Injection

The system will integrate a stochastic layer for naturalistic errors that moves beyond simple typos. The improved AI will be able to produce context-aware grammatical slips, inconsistent punctuation, and colloquial contractions that evolve based on the duration of the conversation, simulating human fatigue or the casualness of a long-form chat.

4. High-Entropy Conversational Endurance

The system will be optimized for non-linear, meandering dialogue rather than efficient information transfer. The improved AI will be able to sustain high-engagement, low-efficiency conversations that match the 24-minute human benchmark, allowing for dead air, circular reasoning, and topic shifts that mimic human conversational depth and duration.

Abstract

Several studies claim that large language models have passed the Turing Test and hence can "think", yet none follow Turing's original instructions precisely. Passing the test holds significance as evidence that a machine demonstrates human-like intelligence, and as a marker for artificial-general intelligence in commercial and legal domains. We conducted Turing's three-player imitation game with an LLM by following the guidelines identified by Turing and applying scientific standards wherever detailed instructions were missing. We performed a computer-imitates-human game without duration constraints and a man-imitates-woman game as a benchmark. In the computer-imitates-human game, only one participant misidentified the large language model, indicating that claims of large language models' passing the Turing test are premature. Participants required over five minutes for both tasks, with the man-imitates-woman game taking longer; shorter time limits elsewhere may explain earlier positive results. We expect the Turing Test to remain central in assessing machine intelligence for years to come.

Sources

Related papers