Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration

arXiv:2608.05741 · cs.CL, cs.AI · Submitted 2026-08-06 · Read on arXiv

Institute of Information Engineering, Chinese Academy of Sciences · School of Cyber Security, University of Chinese Academy of Sciences · College of Computer Science and Technology, Zhejiang University · Institute of Computing Technology, Chinese Academy of Sciences

cs.CL, cs.AI

Submitted: 2026-08-06

Updated: 2026-09-26

Comments: 17 pages, 7 figures

Code: https://github.com/meta-llama/llama-models

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: The paper addresses the growing challenge of detecting machine-generated text.

Terminology

Summary

The paper addresses the growing challenge of detecting machine-generated text. As the authors state, Large language models (LLMs) can generate fluent and convincing text at scale, creating growing risks for misinformation dissemination, educational misuse, and platform governance. They note that prior work has shown humans themselves often struggle to reliably distinguish model-generated text from human writing, making robust detection increasingly essential for the safe and trustworthy deployment of LLM systems.

The authors identify a limitation in existing zero-shot detectors: "Recent zero-shot detectors mainly exploit probability-based statistical discrepancies, but they do not explicitly account for the training process of LLMs, which leaves a distinct generation mechanism insufficiently modeled and limits detection robustness. Building on the insight from IRM that instruction tuning leaves detectable traces in generated text, the paper asks a natural follow-up question: beyond changing token-level probabilities, does instruction tuning leave behind a more persistent signal that fundamentally distinguishes machine-generated text from human writing?"

The central hypothesis is that "this persistent signal comes from the latent assistant-response context introduced by instruction tuning. Unlike human-written text, machine-generated text is typically produced as a response under a system prompt, a user request, or an assistant role. Although this original prompt is removed, its influence is not fully erased: the generated text still implicitly 'remembers' that it was written as a response. Crucially, this hidden dependency can be reactivated without knowing the true prompt. In particular, simply prepending a unified generic assistant-style prefix makes machine-generated text align more naturally with the induced conditioning context, whereas human-written text exhibits a much weaker response to the same prefix."

EchoPrompt operates in three steps:

Step 1: Assistant-context restoration. A unified task-agnostic prefix cg is prepended to the input text to construct a restored sequence [cg; X]. The chosen prefix is: "You are a helpful, versatile, and intelligent AI assistant. Below is the content you generated in response to a user's request, which acts as either a coherent continuation, a topic-specific article, or a detailed answer to a question: This prefix is intentionally task-agnostic... it restores only the coarse global condition that the following sequence should be interpreted as an assistant-style response."

Step 2: Context-calibrated comparative scoring. The method computes token-level log-likelihoods under two asymmetric conditions: the instruction-tuned proxy model evaluates the restored sequence [cg; X], while the corresponding base model evaluates the original text X to calibrate ordinary linguistic predictability. The EchoPrompt score is formally defined as:

ScoreEchoPrompt (X; cg) = (1/(n−1)) Σ t=2 n [log Pinst(xt cg, x<t) − log Pbase(xt x<t)]

The authors explain that "The first term measures how naturally the passage is supported under assistant-style contextual conditioning, while the second term provides a calibration baseline for its local linguistic predictability. Their difference suppresses fluency effects shared by both human and machine text, and highlights the additional advantage that machine-generated passages receive when evaluated under the restored assistant-style condition."

Step 3: Threshold-based detection. The final classification compares the score against a threshold τ, classifying passages above the threshold as AI-generated and those below as human-written. A higher EchoPrompt score indicates stronger latent dependency on the restored assistant-style context, and therefore a higher likelihood of machine generation.

To validate the central intuition, the authors directly measure the likelihood gain induced by prompt injection within the same base model, defined as g(X; cg) = (1/(n−1)) Σ [log P(xt cg, x<t) − log P(xt x<t)]. The results show "a consistent pattern emerges: after the same generic prefix is injected, machine-generated text receives a larger likelihood gain than human-written text. This result indicates that machine text is more naturally compatible with the restored assistant-style condition. The paper emphasizes that the useful signal captured by EchoPrompt is not merely raw fluency, but the extra advantage a passage obtains when evaluated under an assistant-style contextual prompt."

The evaluation uses three public benchmarks: DetectRL (with Multi-Domain, Multi-LLM, and Multi-Attack splits), RealDet, and RAID. Baselines include training-based methods (OpenAI-D, BiScope, R-Detect) and training-free methods (Likelihood, LogRank, Entropy, Fast-DetectGPT, Binoculars, LastDE++, DNA-DetectLLM, IRM). Proxy model families include Qwen2.5-1.5B/3B, Llama-3.2-1B/3B, Llama-3.1-8B, Meta-Llama-3-8B, and Falcon-7B, with paired base/instruct models. Metrics are AUROC and F1 score.

Using the Llama-3-8B proxy family, EchoPrompt achieves the strongest overall performance... improving over the strongest training-free baseline, IRM, by 0.69% AUROC and 2.64% F1 on average. Compared to the best training-based detector OpenAI-D, the gains are much larger, reaching 13.12% AUROC and 17.18% F1 on average. On the distributionally different RealDet and RAID benchmarks, EchoPrompt improves over the second-best method by 4.33% F1 on RealDet and by 1.64% AUROC / 1.38% F1 on RAID. The authors attribute this advantage to the design: "By restoring a task-agnostic assistant-response context, EchoPrompt exposes this latent generation dependency. The base-model comparison further filters out ordinary linguistic predictability, leaving a cleaner signal of assistant-style contextual compatibility."

EchoPrompt obtains the best AUROC and F1 in four out of five attack groups and achieves the strongest average attack performance, improving over IRM by 0.24% AUROC and 1.37% F1 on average. The gains are most pronounced under direct prompting and perturbation, where EchoPrompt improves F1 over the second-best method by 2.74% and 2.20%, respectively. Under paraphrasing, it ranks second but remains very close to IRM, trailing by only 0.87% AUROC and 1.10% F1 while still achieving a high F1 score of 94.53%. The robustness stems from the fact that EchoPrompt focuses on whether the text retains generation-style dependency rather than on isolated token statistics.

Prompt choice. The full prefix outperforms the empty-prompt setting, improving AUROC by 14.73% and 12.57% on Qwen2.5-3B, and by 5.63% and 5.33% on Llama-3.1-8B. The component-level analysis reveals that context clause A, which frames the passage as content generated in response to a user's request, provides the strongest individual contribution, while removing context clause A causes the largest drops on Qwen2.5-3B, with AUROC decreasing by 11.59% and 11.88%. This confirms that the core signal of EchoPrompt comes from restoring the missing prompt–response relation.

Proxy models. Averaged over all proxy settings, EchoPrompt achieves 87.06% AUROC and 83.03% F1, outperforming the strongest competing average baseline, DNA-DetectLLM, by 2.70% AUROC and 3.23% F1. On Llama-family proxies it reaches 93.50% AUROC and 88.25% F1 on average, exceeding IRM by 1.80% AUROC and 2.63% F1.

Text lengths. "Even in the shortest 1–40 word bin, EchoPrompt remains competitive, with an average AUROC of 75.8% across the three Llama-family proxies. As length increases, its performance rises rapidly to 93.1% in the 81–120 bin and 98.9% in the 321–360 bin."

The runtime analysis shows EchoPrompt falls in the middle range... it remains efficient relative to heavier baselines and provides a favorable efficiency–accuracy trade-off. The paper reports an inference latency of less than 0.26 seconds per sample (0.254s, exactly 1.00x relative to its own baseline, compared to 1.529s for LastDE++ at 6.03x and 0.444s for DNA-DetectLLM at 1.75x).

The authors acknowledge that "Like other zero-shot detectors, EchoPrompt still depends on the choice of proxy family. In addition, the present prompt study shows that adding semantic context is useful, but it does not establish that the current prompt is globally optimal. For broader impacts, they note the potential to help with misinformation mitigation, educational integrity, authorship transparency, and platform governance, but caution that False positives may wrongly label human-written text as machine-generated, and false negatives may miss generated content... Therefore, EchoPrompt should be used as an auxiliary signal, not as definitive evidence of authorship."

The paper concludes that "By combining generic prefix restoration with calibrated likelihood comparison between base and instruction-tuned models, EchoPrompt captures contextual traces left by the generation process... These results highlight latent prompt dependency as an effective signal for zero-shot machine-generated text detection."

Improvements for AI systems

Improvements to AI systems based on EchoPrompt:

  1. Add a training-free, out-of-the-box AI-text detector to content moderation pipelines.

The system can immediately classify incoming text as human‑ or machine‑written across domains and models, with no fine‑tuning needed. It works by prepending a task‑agnostic assistant prefix, scoring with an instruction‑tuned proxy model, and subtracting the base model’s likelihood for calibration.

  1. Make detection robust to paraphrasing and perturbation attacks.

Because EchoPrompt measures the latent assistant‑response dependency rather than surface token statistics, the improved system can still flag AI‑generated text after paraphrasing, word‑level perturbations, or direct prompt injection. This yields consistently high AUROC and F1 under adversarial conditions.

  1. Dramatically reduce false positives on fluent human writing.

By contrasting the instruct‑model likelihood (with restored assistant context) against the base‑model likelihood (ordinary fluency), the system cancels out generic linguistic predictability. It therefore avoids penalizing eloquent human prose and only flags text that truly benefits from the assistant‑style conditioning.

  1. Deploy lightweight, real‑time detectors on edge devices or APIs.

Using small paired proxy models (e.g., 1B–3B parameters), EchoPrompt achieves near‑SOTA accuracy while running under 0.26 seconds per sample. The improved system can be integrated into browser extensions, LMS platforms, or social‑media moderation tools without large GPU infrastructure.

  1. Make detection thresholds adaptive for different safety‑vs‑recall requirements.

The EchoPrompt score is continuous and interpretable—it reflects the likelihood gain from restoring the assistant context. The improved system can adjust its decision threshold based on the application (e.g., high precision for legal evidence, high recall for misinformation filtering), using a small validation set or a desired operating point.

  1. Generalize to new LLMs and unseen domains without retraining.

Since the method is zero‑shot and only requires any paired base/instruct model family, the improved system can detect outputs from future or proprietary LLMs as soon as they are released, without collecting new training data for each model.

  1. Provide a verifiable, interpretable signal for AI‑authorship auditing.

The per‑text likelihood gain directly quantifies how much the text “remembers” being an assistant response. The improved system can output this score alongside the classification, enabling audits, reporting, and human review with clear evidence of why a text was flagged.

  1. Enable proactive governance and platform safety.

With higher robustness and lower false positives, the improved system can be deployed for misinformation mitigation, educational integrity checks, authorship transparency, and enforcement of content policies—while being transparent about its role as an auxiliary signal rather than definitive proof of authorship.

Sources

Related papers