The Challenge of Identifying the Origin of Black-Box Large Language Models
cs.CR, cs.LG
Submitted: 2025-03-06
Updated: 2026-09-21
Comments: To Appear in the Findings of the Association for Computational Linguistics: EMNLP 2026
Code: https://github.com/OpenBMB/MiniCPM-V
License: http://creativecommons.org/licenses/by/4.0/
The gist: The tremendous commercial potential of large language models (LLMs) has heightened concerns over their unauthorized use.
Terminology
Abstract
The tremendous commercial potential of large language models (LLMs) has heightened concerns over their unauthorized use. To address this, we focus on the task of identifying the origin of black-box LLMs. We further propose PlugAE, an effective and efficient identification method that proactively leverages LLM-specific adversarial embeddings and allows users to customize copyright tokens on a targeted query set. Extensive experiments demonstrate that PlugAE outperforms both state-of-the-art model watermarking and fingerprinting methods in accuracy and robustness. We further analyze its stealthiness and reliability from three complementary perspectives and conduct ablation studies under various configurations, confirming its practicality for real-world misuse detection.
Sources
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2
- Mistral 7B
- ProFLingo: A Fingerprinting-based Intellectual Property Protection Scheme for Large Language Models
- OLMo: Accelerating the Science of Language Models
- Watermarking Pre-trained Language Models with Backdooring
- Your Large Language Models Are Leaving Fingerprints
- Gemma: Open Models Based on Gemini Research and Technology
- CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
- Instructional Fingerprinting of Large Language Models
- SOS! Soft Prompt Attack Against Open-Source Large Language Models
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- LLaMA: Open and Efficient Foundation Language Models
- HuRef: HUman-REadable Fingerprint for Large Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- CogVLM: Visual Expert for Pretrained Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs