Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot
cs.CR, cs.AI
Submitted: 2026-06-14
Updated: 2026-08-29
Terminology
Sources
- Stealing Part of a Production Language Model
- MISLEADER: Defending against Model Extraction with Ensembles of Distilled Models
- Increasing the Cost of Model Extraction with Calibrated Proof of Work
- GPT-4 Technical Report
- The Llama 3 Herd of Models
- Model Leeching: An Extraction Attack Targeting LLMs
- GPT-4o System Card
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing Attacks
- Thieves on Sesame Street! Model Extraction of BERT-based APIs
- HoneyGPT: Breaking the Trilemma in Terminal Honeypots with Large Language Model
- GreaseLM: Graph REASoning Enhanced Language Models for Question Answering
- NSmark: Null Space Based Black-box Watermarking Defense Framework for Language Models
- LLMPot: Dynamically Configured LLM-based Honeypot for Industrial Protocol and Physical Process Emulation
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs