An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic
cs.CR, cs.CL
Submitted: 2026-06-04
Updated: 2026-08-29
Comments: Accepted to EMNLP 2026 Main Conference. Code available at https://github.com/LabRAI/mmd-llm-mea-detection
Code: https://github.com/LabRAI/mmd-llm-mea-detection
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) are increasingly deployed through hosted APIs, making model extraction a practical threat to model ownership and service security.
Terminology
Abstract
Large language models (LLMs) are increasingly deployed through hosted APIs, making model extraction a practical threat to model ownership and service security. Individual extraction queries often resemble benign requests, while existing evaluations often focus on single-query anomaly scoring or pure benign-versus-attacker user settings. We formulate model extraction monitoring as benign-calibrated traffic-window distribution testing: embed incoming queries into a semantic space and test whether their aggregate distribution deviates from historical benign traffic. We instantiate this formulation with maximum mean discrepancy (MMD), using only benign-vs-benign comparisons to set the decision threshold. We evaluate on fourteen attacker-normal query pairs from four extraction scenarios and compare with adapted PRADA, SEAT, CAP, DATE, marginal Mahalanobis, and pseudo-class energy baselines. Across three random seeds, MMD achieves 0.3% benign FPR, 100.0% pure-attacker TPR, 90.5% average TPR over attacker fractions, and 95.1% balanced accuracy. These results show that benign-calibrated distribution testing is a strong empirical baseline for model extraction detection in both user-level and mixed multi-user LLM API traffic.
Sources
- Model Leeching: An Extraction Attack Targeting LLMs
- On the Opportunities and Risks of Foundation Models
- Distilling the Knowledge in a Neural Network
- Pointer Sentinel Mixture Models
- A Survey on Model Extraction Attacks and Defenses for Large Language Models
- WildChat: 1M ChatGPT Interaction Logs in the Wild
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs