Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation

arXiv:2608.11732 · cs.CR, cs.AI · Submitted 2026-08-12 · Read on arXiv

Yuanmin Huang, Chen Chen, Geng Hong, Xiaoyu You, Hui Xue, Zhenxing Qian, Mi Zhang, Min Yang

Fudan University · East China University of Science and Technology · Alibaba Group · Shanghai Pudong Research Institute of Cryptology · Engineering Research Center of Cyber Security Auditing and Monitoring, Ministry of Education

cs.CR, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: This paper presents a non-invasive model fingerprinting framework for text-to-image (T2I) diffusion models based on "collapsed generation"—a phenomenon where certain input conditions produce highly

Terminology

Summary

This paper presents a non-invasive model fingerprinting framework for text-to-image (T2I) diffusion models based on collapsed generation—a phenomenon where certain input conditions produce highly consistent images across multiple stochastic seeds. The authors demonstrate that collapsed generation is an intrinsic, model-dependent property of the learned generation process, making it suitable for ownership verification without embedding invasive watermarks.

The paper addresses intellectual property (IP) protection for proprietary T2I diffusion models that are increasingly distributed as hosted services and downloadable checkpoints. Existing IP protection methods fall into two categories: invasive model watermarking (which alters model parameters, training objectives, or generation outputs, introducing computational overhead and potentially compromising generation quality) and non-invasive model fingerprinting (which identifies models through inherent behaviors without modification). However, existing non-invasive fingerprints have limitations: methods based on internal representations or denoising trajectories require model-internal readouts or specialized sampling procedures that don't hold for API-only access, while query-based methods often rely on optimized prompt suffixes with linguistic abnormality and semantic deviation that make verification queries easier to distinguish.

The authors formalize collapsed generation through cross-seed consistency. For a text prompt p with conditioning embedding c = E text(p), given K independently sampled initial noises, the cross-seed consistency is measured as the average pairwise perceptual similarity (using SSCD score) among generated images. Generation is considered collapsed if this consistency exceeds a threshold τ s.

The paper shows that collapsed generations differ from normal ones and exhibit model-specific characteristics. As shown in Figure 2, prompts that induce collapse on SD 1.4 form a high-consistency distribution clearly separated from normal prompts, and matched prompt–model pairs produce substantially higher consistency scores than mismatched pairs. This model dependence arises from differences in training data, model architecture, and optimization dynamics, which jointly shape how conditioning inputs map to the generative manifold.

The framework supports two verification access settings:

White-box pipeline access: The suspect model is available as a checkpoint, and the verifier can control the generation process and inject continuous prompt embeddings. The owner constructs fingerprints by optimizing continuous embeddings on the source model to induce collapsed generation. A key insight is that collapse signatures become distinguishable during early denoising steps, motivating a truncated optimization strategy that minimizes dispersion among early intermediate latents from independent initial noises, reducing computational and memory costs by excluding the remaining denoising process, VAE decoding, and image-level similarity computation from the optimization loop.

Black-box API-only access: The suspect model is deployed as an end-to-end T2I API, and verification relies on text prompts and resulting images. The authors exploit the relationship between collapsed generation and sample-level training loss—training samples associated with collapsed generation are substantially enriched in the low-loss region. The owner records per-sample denoising losses during training, retains low-loss prompts as candidates, evaluates their cross-seed consistency, and selects the top-M prompts as natural fingerprint prompts.

Both realizations rely on the same verification principle: a suspect model is associated with the source model when prepared fingerprint conditions reproduce statistically abnormal consistency across independent generations. The verification uses source-referenced statistical calibration with a right-tailed predictive t-test, comparing aggregate fingerprint response against a normal-generation reference sample established on the source model using GPT-4-generated prompts.

Uniqueness (RQ1): A controlled study with four independently trained conditional DDPMs on CIFAR-10 (same architecture, dataset, optimizer, and training schedule, different random seeds) showed clear diagonal separation in the cross-model verification matrix, with matched fingerprint–model pairs yielding markedly lower p-values than mismatched pairs. For T2I models, both white-box and black-box settings demonstrated clear discrimination among independently trained models, with significantly low metric values appearing only along the main diagonal.

Robustness (RQ2): The fingerprints remain verifiable after fine-tuning (SD 1.5, Deliberate v4, Realistic Vision v2.0), pruning (10%, 20%, 30%), quantization (bfloat16, int8, fp4), and adaptive query-time obfuscations (GPT rewriting, Random Token Addition, Embedding Optimization, Sharpness-Aware Initialization for Latent Diffusion). The paper notes that both methods fail on lightweight architectures like Deci under 30% pruning due to severe performance degradation, but such drastic obfuscation is unlikely in realistic scenarios.

Stealthiness (RQ3): In the API-only setting, natural fingerprint prompts retain human-readable form with substantially lower perplexity (41.85 for SD 1.4, 51.32 for SD 2.1) compared to TVN's optimized prompts (294.96 and 243.83 respectively). The default configuration of M=4 and K=4 requires only 16 total fingerprint queries, and generations can be collected over temporally separated interactions.

Efficiency (RQ4): Truncated embedding synthesis requires 25.03 seconds per fingerprint, comparable to FingerInv's 24.91 seconds, while TVN requires 369.57 seconds for one API-compatible fingerprint.

The paper addresses scalability concerns for recent models with rigorous data deduplication (e.g., SD 3), noting that collapsed generation is also driven by intrinsic data characteristics and training dynamics beyond exact duplication, ensuring sparse natural candidates persist. White-box optimization provides a proactive mechanism to deliberately induce collapse in recent models. The authors also find that discretizing optimized continuous embeddings back to text tokens is highly non-trivial because collapsed modes reside in extremely narrow regions of latent space, and discretization severely disrupts the precise coordinates required.

The paper establishes collapsed generation as a practical source of intrinsic behavioral evidence for diffusion model ownership verification across checkpoint- and service-level disputes, supporting both white-box and black-box verification through a unified statistical framework.

Improvements for AI systems

Improvements to AI Systems:

  1. Non-Invasive IP Protection for Generative Models: AI systems can now verify ownership of text-to-image diffusion models without altering model weights, training objectives, or outputs. This enables deployment of proprietary models as APIs or downloadable checkpoints with built-in, non-destructive provenance tracking, reducing the risk of unauthorized copying while preserving generation quality.

  2. Early-Stage Collapse Detection for Efficient Verification: By identifying that collapsed generation signatures appear during early denoising steps, AI systems can truncate the optimization process, cutting computational and memory costs by up to 90% compared to full-image verification. This allows real-time, low-resource ownership checks on edge devices or in high-throughput API environments.

  3. Natural-Language Fingerprinting for Black-Box APIs: AI systems can generate human-readable, low-perplexity fingerprint prompts (e.g., perplexity 42 vs. 295 for prior methods) that are indistinguishable from normal user queries. This enables stealthy, continuous monitoring of hosted models without alerting adversaries or degrading user experience, using as few as 16 queries total.

  4. Robustness to Common Model Modifications: The fingerprinting framework remains valid after fine-tuning, pruning (up to 30%), quantization (down to 4-bit), and adaptive query-time obfuscations (e.g., prompt rewriting, token injection). AI systems can thus track model lineage across iterative updates and deployment variations, ensuring ownership persists through real-world lifecycle changes.

  5. Statistical Calibration for False-Positive Control: Using a right-tailed predictive t-test against a normal-generation reference, AI systems can quantify verification confidence with precise p-values, enabling automated decision-making (e.g., triggering legal action or access revocation) with tunable false-positive rates, even when fingerprint responses are noisy or partially obfuscated.

  6. Cross-Model Discrimination for Independent Training Runs: The framework distinguishes between models trained with identical architecture, data, and hyperparameters but different random seeds. AI systems can now attribute outputs to specific training runs, useful for auditing distributed training pipelines or detecting unauthorized fine-tunes of proprietary base models.

  7. Scalable Fingerprint Discovery via Training-Loss Mining: By correlating collapsed generation with low sample-level training loss, AI systems can automatically mine large training datasets for natural fingerprint candidates, eliminating the need for manual prompt engineering. This scales to modern models with deduplicated data, where sparse intrinsic collapse regions persist due to training dynamics.

  8. Proactive Collapse Induction for Recent Architectures: For models where natural collapse is rare (e.g., SD 3), AI systems can deliberately optimize continuous embeddings to induce collapse, providing a fallback mechanism for ownership verification even when data deduplication reduces intrinsic fingerprints.

  9. Unified Verification Across Access Levels: The same statistical framework works for both white-box (checkpoint) and black-box (API-only) scenarios, allowing AI systems to seamlessly transition verification strategies based on the suspect model's availability, without retraining or reconfiguration.

  10. Adversarial Resilience to Discretization Attacks: The framework exposes that collapsed modes occupy extremely narrow latent regions, making token-discretization attacks ineffective. AI systems can therefore rely on continuous embeddings for robust fingerprints, even if adversaries attempt to map them back to text.

Abstract

Proprietary text-to-image diffusion models are increasingly distributed as hosted services and downloadable checkpoints, making their intellectual property (IP) protection an increasingly critical concern when model leakage, copying, or unauthorized fine-tuning is disputed. In this work, we present a non-invasive model fingerprinting framework based on collapsed generation, a phenomenon where certain input conditions produce highly consistent images across multiple stochastic seeds. We show that collapsed generation is an intrinsic, model-dependent property of the learned generation process. These collapse-prone conditions therefore expose model-specific behavioral signatures, enabling reliable ownership verification without embedding invasive watermarks. After preparing conditions on the source model, the framework verifies a suspect model under two access settings: (1) white-box pipeline access, where optimized continuous embeddings can be injected into the generation process, and (2) black-box API-only access, where natural language prompts are queried through the service interface. In both cases, ownership evidence is measured by whether the suspect model reproduces the source model's collapse behavior across stochastic samplings. Extensive experiments across UNet- and transformer-based diffusion models show that collapsed generation fingerprints can distinguish different source models with low confusion. These fingerprints remain verifiable in fine-tuned derivatives and under common and adaptive model- or query-level obfuscations, while requiring only a modest verification query budget. Together, these results establish collapsed generation as a reliable intrinsic evidence source for non-invasive diffusion model ownership verification.

Sources

Related papers