Fingerprinting Multimodal Large Language Models
cs.CR, cs.AI
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: 10 pages, 3 figures. Accepted to ACM Multimedia 2026 (MM '26) as an oral presentation
License: http://creativecommons.org/licenses/by/4.0/
The gist: While multimodal large language models (MLLMs) enable a wide range of image-text reasoning tasks, recent incidents indicate that they are vulnerable to illicit deployment and unauthorized
Terminology
Abstract
While multimodal large language models (MLLMs) enable a wide range of image-text reasoning tasks, recent incidents indicate that they are vulnerable to illicit deployment and unauthorized distillation. Existing solutions for model provenance are typically confounded by shared language backbones in MLLMs and struggle to detect violations of distillation. To bridge this gap and safeguard model ownership, we present the first study on multimodal model fingerprinting. Inspired by recent findings that self-attention acts as a low-pass filter and that its low-frequency components are informative, we develop AttnPrint for white-box provenance. Specifically, we extract cross-modal attention distributions and isolate their low-frequency components to serve as model fingerprints. To facilitate black-box auditing, we further introduce DistillTrace, which employs hypothesis testing of MLLM outputs to identify potential model infringement. We conduct extensive experiments on 154 model instances across 19 multimodal architectures. Notably, AttnPrint achieves strong derivative-model detection performance while remaining robust to five downstream modification techniques. DistillTrace also provides evidence of distillation relationships under three parameter-independent techniques.
Sources
- Pixtral 12B
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Streamlining Redundant Layers to Compress Large Language Models
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Gemma 3 Technical Report
- Distilling the Knowledge in a Neural Network
- LoRA: Low-Rank Adaptation of Large Language Models
- GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning
- Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
- Building and better understanding vision-language models: insights and future directions
- What matters when building vision-language models?
- LLaVA-OneVision: Easy Visual Task Transfer
- Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
- ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
- Connecting Vision and Language with Localized Narratives
- Detecting Contextual Hallucinations in LLMs with Frequency-Aware Attention
- SoK: Large Language Model Copyright Auditing via Fingerprinting
- Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation
- A Simple and Effective Pruning Approach for Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs