MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection

arXiv:2608.10459 · cs.CL, cs.AI · Submitted 2026-08-11 · Read on arXiv

Jinmo Han, Jimin Hong, Chanyeong Moon, Ju Yeon Kang, Seonuk Kim, Nam Soo Kim

Seoul National University

cs.CL, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: MD-ProTector is an input-only encoder detector for LLM-generated text detection that represents each class (human-written and machine-generated) with multiple trainable reference vectors, called

Terminology

Summary

MD-ProTector is an input-only encoder detector for LLM-generated text detection that represents each class (human-written and machine-generated) with multiple trainable reference vectors, called prototypes, in the encoder embedding space. The paper states: "We propose MD-ProTector, which represents each class with multiple trainable reference vectors in the encoder embedding space, referred to as prototypes. These prototypes provide separate decision boundaries for different groups of texts within the same class."

The key challenge addressed is that adding multiple prototypes alone does not determine which variation each prototype should represent. The paper explains: "However, the number of prototypes alone does not determine which pattern of within-class variation each prototype should represent. Without a prototype-specific objective, different prototypes may remain redundant or capture overlapping patterns."

MD-ProTector addresses this with a Prototype Positioning loss, which separates class-level structure from the within-class variation that differentiates individual prototypes. The method separates the direction shared within each class from the variation that distinguishes groups of samples within that class. The paper describes: "MD-ProTector separates the direction shared within each class from the variation that distinguishes groups of samples within that class. Prototype Positioning loss uses this variation to give different prototypes distinct roles, while complementary objectives keep each prototype aligned with the corresponding human or machine class."

The training objectives include:

  • Prototype-to-Class Loss: aligns each prototype with the hub of its own class

  • Sample-to-Prototype Loss: associates each sample with its ground-truth prototype bank while separating it from the opposite-class bank

  • Prototype Positioning Loss: a softmax cross-entropy over the prototype residual vectors, with a separate data-derived target constructed for each prototype

The inference procedure is described as: Given an input text x, we compute z = norm(fθ(x)) and score each class by its most similar prototype with the detection score S(z) = s1(z) − s0(z), where s0(z) and s1(z) are the maximum similarities to the machine and human prototypes, respectively.

Prototypes are initialized using K-Means separately within each class: Before training, we extract embeddings from the training set and apply K-Means separately within each class. The resulting centroids initialize Pc = pc,r R r=1.

Evaluation results show: MD-ProTector achieves the highest AvgRec on MAGE CDCM and RAID and the highest AUROC and lowest FPR95 on RAID among the compared encoder-based methods. Specifically:

  • On MAGE CDCM: MD-ProTector obtains the highest AvgRec on both MAGE CDCM and RAID. On MAGE CDCM, it reaches 95.14 AvgRec with balanced recalls of 95.81 for human and 94.47 for machine text.

  • On RAID: MD-ProTector achieves the highest AvgRec (88.18), HumanRec (82.52), and AUROC (95.41), together with the lowest FPR95 (27.78).

  • On M4: It obtains the second-highest AvgRec of 86.03, improving HumanRec to 76.87 compared with 54.56 for Binary CE, 72.37 for SupCon, and 63.13 for DSVDD, while retaining 95.20 MachineRec.

Ablation studies show: Removing LPP lowers AvgRec from 95.14 to 94.78. Replacing Prototype Positioning with simple prototype repulsion yields 94.55, while positioning prototypes without removing the class-hub direction yields 94.33. K-Means initialization improves AvgRec from 94.50 to 95.14 relative to random initialization. The optimal number of prototypes is R = 8, and the optimal temperature is τ = 0.15.

The paper concludes: "By separating class-level alignment from prototype-specific residual positioning, MD-ProTector organizes multiple prototypes using the variation observed within each class while retaining direct prototype-based inference. Across five controlled settings, the method achieves the highest AvgRec on MAGE CDCM and RAID. On RAID, it also attains the highest AUROC and lowest FPR95 among the compared methods."

Improvements for AI systems

Improvements to AI Systems:

  1. Enhanced Few-Shot and Imbalanced-Class Text Detection: By using multiple trainable prototypes per class with the Prototype Positioning loss, an AI system can better model within-class variation (e.g., different writing styles, topics, or domains) without collapsing into a single centroid. This improves detection accuracy for underrepresented or heterogeneous text groups, reducing false negatives for human-written text in imbalanced datasets (e.g., M4 HumanRec improved from 54.56 to 76.87).

  2. Interpretable Decision Boundaries for Text Classification: The system can assign each input to the most similar prototype, providing a human-understandable explanation (e.g., this text matches the 'formal academic' human prototype or this matches the 'casual chat' machine prototype). This enables users to see which stylistic variation triggered the classification, aiding transparency in AI content moderation or plagiarism detection.

  3. Robust Out-of-Distribution Detection: The prototype-based inference with class-hub separation allows the system to compute residual vectors that capture within-class variation. This makes the detector more sensitive to novel or adversarial text patterns that do not align with any prototype, lowering FPR95 (27.78 on RAID) and improving AUROC (95.41) for distinguishing AI-generated from human text under distribution shift.

  4. Efficient Adaptation to New Domains or Languages: The K-Means initialization and prototype-specific objectives allow the system to quickly re-train or fine-tune prototypes on new data without full model retraining. An AI system can incrementally add prototypes for emerging writing styles (e.g., new LLM versions or new languages) while preserving existing class structure, reducing computational cost and catastrophic forgetting.

  5. Controllable Precision-Recall Trade-off: Since inference uses per-class maximum similarity, the system can adjust detection thresholds per prototype or per class. This allows deployment in high-recall settings (e.g., flagging all possible AI text for review) or high-precision settings (e.g., only flagging confident matches), making it adaptable to different regulatory or user needs.

What the Improved AI System Can Do:

  • Detect AI-generated text with state-of-the-art average recall (95.14 on MAGE CDCM, 88.18 on RAID) while maintaining balanced performance across human and machine classes.

  • Provide granular, explainable classifications by identifying which specific prototype (e.g., news-style human vs. creative-writing machine) matches the input.

  • Operate reliably on diverse, mixed-domain corpora (RAID, M4) with low false-positive rates, suitable for real-world content moderation, academic integrity checks, and social media filtering.

  • Rapidly adapt to new LLM outputs or writing styles by updating prototype banks, without sacrificing existing class knowledge.

Abstract

As LLM-generated content becomes more sophisticated, detection systems for distinguishing those texts from human-written text must operate at scale while handling diverse writing styles, domains, languages, and generator models. Input-only encoder detectors are suitable for practical deployment setting, but standard binary classification supplies only the class label and does not explicitly organize the substantial variation within either class. We propose MD-ProTector, which represents each class with multiple trainable reference vectors in the encoder embedding space, referred to as prototypes. These prototypes provide separate decision boundaries for different groups of texts within the same class. However, adding multiple prototypes alone does not determine which variation each prototype should represent. MD-ProTector addresses this problem with Prototype Positioning loss, which separates class-level structure from the within-class variation that differentiates individual prototypes. Evaluated across five settings from three large-scale benchmarks covering domain, generator, language, and adversarial variation, MD-ProTector achieves the highest AvgRec on MAGE CDCM and RAID and the highest AUROC and lowest FPR95 on RAID among the compared encoder-based methods.

Sources

Related papers