Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs

arXiv:2603.02353 · cs.CL · Submitted 2026-03-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs".

Jane: The paper was written by Jiangang Hao from ETS Research Institute.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We are looking at a fascinating new paper today called "Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs" by Jiangang Hao from the ETS Research Institute.

Jane: It sounds like a mouthful, Tom, but it's really addressing a question every teacher is asking right now.

Tom: Are you asking if the student actually wrote that essay or if a machine did it for them?

Jane: Exactly, and Hao is looking at this specifically within standardized tests where the rules are much stricter than a regular classroom.

Lu: This is such a massive shift in how we define authorship in the digital age!

Meng: I want to know if these detection tools can actually be used in a real testing center without breaking everything.

Lalam: It touches on the very heart of how we will value human thought as these models become part of our daily lives.

Tom: That's the tension the title is hinting at when it mentions "Responsible Use."

Jane: It's a warning that we can't just throw these detectors at students and hope for the best.

Lu: We could see a future where the detector itself becomes a part of the creative dialogue!

Meng: If the detector is wrong, a student could lose a huge opportunity, so we need to know how reliable these things are.

Lalam: We have to ensure that the technology helps us see the person behind the words rather than just flagging them as errors.

Tom: That leads us directly into how they actually tested these ideas in the research.

Summary: Tom: To figure out if detectors work, Hao tested them against a huge range of models, including GPT-four GPT-4o, and even GPT-five.

Jane: They used something called "perplexity" to see how predictable the writing is, using GPT-two as a benchmark to measure that.

Tom: So, if the text is too predictable, the detector starts getting suspicious?

Jane: That's the idea, and they measured success using a score called AUC to see how well they could separate human writing from machine writing.

Lu: I was stunned by the results showing that GPT-five and GPT-o4-mini form their own little group that's totally different from the older models!

Meng: That's a huge practical problem because a detector trained on GPT-four might completely miss a GPT-five essay.

Lalam: It's like the language is evolving into entirely new species that our old tools don't recognize.

Tom: The paper mentions a "GPT-all" approach where they train the detector on every model at once to fix that.

Jane: That seems much more robust than just focusing on one specific version of a model.

Meng: It makes sense to build a net that catches everything instead of just one type of fish.

Lu: We are seeing the birth of a universal linguistic signature for these machines!

Lalam: Even with that, the way these models cluster together shows how much they are starting to share a single voice.

Tom: But even a "universal" detector has some serious blind spots.

Improvements: Tom: The research shows that looking at the text alone isn't a silver bullet for catching AI.

Jane: They suggest looking at the writing process itself, like tracking keystrokes and how much time a student spends on a sentence.

Tom: So, if someone just copies and pastes a whole essay, the lack of typing patterns would give them away?

Jane: Yes, because humans have these natural, irregular rhythms when they type and revise their work.

Meng: That sounds like a great way to catch people, but what about the students who use AI to help them brainstorm and then type it out themselves?

Lu: We could use those behavioral traces to create a beautiful map of how a human mind actually constructs an idea!

Lalam: Perhaps we should start valuing the struggle of the writing process as much as the final product.

Tom: That's a tough one for current grading systems that only care about the finished essay.

Jane: And we can't forget that short responses are much harder to detect because there isn't enough data to find a pattern.

Meng: I also worry about the "hybrid" problem where a person and an AI have worked on the same text together.

Lu: That's where the most interesting kind of human-machine collaboration will happen!

Lalam: We might need to move toward assessments that celebrate how we use these tools to expand our own reasoning.

Tom: It sounds like the way we test writing will have to change completely.

Conclusion: Tom: We've covered a lot, from the technicalities of perplexity to the way keystrokes might save academic integrity.

Jane: It's clear that while detectors are getting better, they aren't a perfect solution on their own.

Tom: We need to combine text analysis with process data and use it very carefully in high-stakes situations.

Lu: I'm just so excited to see how this pushes us to define what makes human expression so unique!

Meng: I'll be watching to see if engineers can build a "GPT-all" detector that actually stays ahead of the next model release.

Lalam: I believe this will lead us to a more honest way of communicating in a world full of synthetic voices.

Jane: Thank you all for joining us to discuss "Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs."

Tom: We'll see you next time for the next big paper!

Jane: Goodbye everyone!

Jiangang Hao

ETS Research Institute

cs.CL

Submitted: 2026-03-04

Updated: 2026-08-21

Importance score: 74/100

The gist: The paper "provides an overview of the current landscape of detectors for AI-generated and AI-assisted essays, along with guidelines for their responsible use" and "presents empirical analyses to

Key concepts

Perplexity
Perplexity is a measure used to determine how predictable a piece of writing is. The research uses it to identify machine-generated text; if the writing is too predictable, the detector becomes suspicious. GPT-two serves as a benchmark to measure this predictability.
AUC
AUC is a score used to measure how well a detector can separate human writing from machine writing. It serves as a metric to evaluate the success and accuracy of detection tools in distinguishing between different types of authorship.
Writing process data
This involves analyzing how a student constructs a text rather than just the final product. By tracking keystrokes and the time spent on sentences, researchers can identify the irregular rhythms of human typing to distinguish real work from AI-generated or copied content.

Terminology

Summary

The paper provides an overview of the current landscape of detectors for AI-generated and AI-assisted essays, along with guidelines for their responsible use and presents empirical analyses to evaluate how well detectors trained on essays from one LLM generalize to identifying essays produced by other LLMs, based on essays generated in response to public GRE writing prompts.

Regarding detection methodologies, the author categorizes approaches into several paradigms. The most common approach treats AI-generated essay detection as a supervised learning problem, utilizing either linguistic and stylistic features or probabilistic indicators such as perplexity or burstiness. This can involve feature-based supervised learning approaches, which tend to exhibit stronger cross-prompt and cross-model generalizability, or end-to-end fine-tuning of pretrained LLMs, such as RoBERTa, which typically achieve slightly higher predictive performance but function largely as black boxes. Another method is Watermarking, which adds a statistical signature, to AI-generated texts, though the watermark signal is highly fragile and requires cooperation from LLM developers. A third direction involves writing process–based features, such as keystroke dynamics, revision histories, and timing information, which can substantially enhance detection performance by capturing the natural, irregular behavioral patterns of human writers. Lastly, Similarity Matching is feasible in standardized writing assessment settings where prompts are fixed, allowing for the comparison of human essays against a large pool of AI-generated essays for each prompt to identify cases where substantial overlap suggests reliance on AI-generated material.

The research on the Generalizability of Detectors Across LLMs evaluated a broader set of more powerful GPT-family models released in 2024 and 2025, including GPT-4, GPT-4o, GPT-o1, GPT-o3-mini, GPT-o4-mini, and GPT-5. Using perplexity-derived features and a Gradient Boosting Machine, the study found that detectors perform very well when evaluated on the same model used during training. A cluster consisting of GPT-4, GPT-4o, GPT-4o-mini, GPT-o1, and GPT-o3-mini consistently achieve AUC values above 0.8 when detecting each other's outputs. However, GPT-o4-mini and GPT-5 show a divergent pattern, as they exhibit poor generalization across the rest of the GPT family. Consequently, a working strategy should be to include training data generated from all LLMs.

Regarding Responsible Use, the paper warns that no AI-generated text detector is perfect and no existing detector performs consistently well across all types of writing tasks or contexts. Specific limitations include: text length also plays a critical role in detector accuracy because performance drops sharply for short responses; most detectors struggle when an essay is jointly produced by humans and AI; and potential bias against certain demographic groups. The paper suggests that institutional consensus is necessary and that detector use should be complemented by evaluation designs that combine take-home assignments with in-class or proctored writing tasks to reduce reliance on detector outputs alone.

For the Future, the paper notes that as AI tools reliably handle surface-level features such as grammar and mechanics, it may be necessary to reconsider the weight these dimensions should carry in scoring. It also emphasizes that detection models must be routinely re-evaluated to monitor performance drift and updated as new LLMs emerge or as patterns in human writing evolve.

Improvements for AI systems

1. Implementation of Cross-Family Multi-Model Training (GPT-all Approach)

  • Improved AI System Capability: The system can detect essays produced by emerging or unseen LLMs (e.g., future iterations of GPT, Gemini, or Claude) by training on a unified, diverse corpus of outputs from all major LLM families rather than model-specific datasets. This mitigates the divergent pattern problem where detectors trained on one model cluster (like GPT-4) fail to identify text from newer, isolated clusters (like GPT-5).

2. Multimodal Authentication via Process-Text Integration

  • Improved AI System Capability: The system can identify hybrid writing (human-edited AI text) and copy-paste plagiarism by correlating text-based statistical signatures with behavioral biometric data. It will specifically detect the absence of natural human compositional traces—such as irregular keystroke dynamics, typing bursts, backtracking, and revision histories—even when the resulting text has been manually modified to bypass linguistic detectors.

3. Deployment of Interpretable Perplexity-Distribution Classifiers

  • Improved AI System Capability: The system replaces black-box end-to-end fine-tuned models with transparent Gradient Boosting Machines (GBM) that utilize high-resolution perplexity-derived features. Instead of a binary score, the system can provide explainable evidence for its judgment by reporting specific sentence-level perplexity distributions (mean, median, and 10th–90th percentiles), allowing human evaluators to diagnose exactly why a text was flagged.

4. Prompt-Constrained Similarity Matching (GPTCollider Integration)

  • Improved AI System Capability: In controlled or standardized assessment environments with fixed prompts, the system can identify AI-reliance by cross-referencing student submissions against a massive, pre-generated pool of LLM-produced essays specific to those exact prompts. This allows for the detection of substantial textual overlap that traditional probabilistic or stylistic detectors might miss.

Sources

Related papers