Can Watermarking Techniques Help Prevent LLM Model Stealing?
Elette Boyle, MohammadTaghi Hajiaghayi, Keivan Rezaei, Suho Shin, Amos Stern
cs.CR
Submitted: 2026-07-12
License: http://creativecommons.org/licenses/by/4.0/
The gist: Model stealing attacks have recently been introduced, enabling the extraction of precise information from black-box commercial language models.
Terminology
Abstract
Model stealing attacks have recently been introduced, enabling the extraction of precise information from black-box commercial language models. In this work, we propose defense methods against a recent attack of and extensions for extracting the hidden layer dimension of production language models. Our methods are inspired by watermarking techniques that perturb the logits layer of these models to prevent such attacks. We provide empirical experiments demonstrating the effectiveness of the proposed defense versus model quality degradation across various configurations, and propose an effective defense against such attacks while preserving model utility.
Sources
- GPT-4 Technical Report
- Polynomial Time Cryptanalytic Extraction of Deep Neural Networks in the Hard-Label Setting
- Stealing Part of a Production Language Model
- The Llama 3 Herd of Models
- Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
- Language Model Inversion
- DeepTextMark: A Deep Learning-Driven Text Watermarking Approach for Identifying Large Language Model Generated Text
- Gemini: A Family of Highly Capable Multimodal Models
- Watermarking Text Generated by Black-Box Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs