Robustness of AI-Art Detectors under Generator Shift
Shivank Singh Thakur, Meien Li, Mark Stamp
San Jose State University
cs.CV, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: To appear as a chapter in the book "Artificial Intelligence for Cyber Defense in Emerging Threats", to be published by Springer by early 2027
Code: https://github.com/pharmapsychotic/clip-interrogator
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper investigates the robustness of AI-art detectors under generator shift, specifically examining whether detectors trained on earlier diffusion generators remain reliable when applied to
Terminology
Summary
This paper investigates the robustness of AI-art detectors under generator shift, specifically examining whether detectors trained on earlier diffusion generators remain reliable when applied to images produced by newer, architecturally different models.
The authors note that Text-to-image generative models have greatly improved in their ability to synthesize high-quality artwork from natural language descriptions
and that modern systems such as DALL·E, Midjourney, and Stable Diffusion now produce images that are often difficult to distinguish from human-created artwork.
This creates a security and trust problem, as AI-generated art can be used for misinformation, fraud, and impersonation.
The central challenge addressed is generator shift: Most existing AI-image detection studies train and evaluate detectors on images from the same generator family or from closely related generators.
The paper models a deployment setting where the defender operates an image-based detector as a screening component at a social media platform, a news verification desk, or a content-moderation pipeline
and has no access to labeled images from future generators.
The authors constructed a prompt-aligned Stable Diffusion 3.5 Medium (SD3.5m) dataset of 10,000 images across ten art styles (Art Nouveau, Baroque, Expressionism, Impressionism, Post-Impressionism, Realism, Renaissance, Romanticism, Surrealism, and Ukiyo-e). The dataset was created through a four-stage pipeline: human artwork selection, reverse prompting via CLIP Interrogator, prompt augmentation and controlled image generation with SD3.5m.
The pipeline used CLIP Interrogator with the BLIP-Large captioning model and the CLIP ViT-L/14 vision encoder
to generate prompts from held-out human artwork samples. Images were generated at a fixed resolution of 768 × 768 pixels using 28 denoising steps and a classifier-free guidance scale of 4.5.
Five detectors were evaluated using a frozen-backbone, linear-probe design
where a pretrained deep network serves as a fixed feature extractor
and a lightweight linear classification head is then trained on top of these frozen features to perform binary human-versus-AI classification.
The models evaluated were:
-
ResNet-18 (feature dimension: 512)
-
ResNet-50 (feature dimension: 2,048)
-
EfficientNet-B0 (feature dimension: 1,280)
-
ConvNeXt-Base (feature dimension: 1,024)
-
CLIP ViT-L/14 (feature dimension: 768)
The in-distribution training data came from AI-ArtBench, which contains approximately 185,000 images across 10 artistic styles
including 60,000 human artwork samples and 125,015 AI-generated images produced by Latent Diffusion (LDM) and Stable Diffusion 2.1 (SD2.1) models.
The training set was balanced through pruning to contain 3 × 4,000 × 10 = 120,000 images with equal representation across all style-source combinations.
All models performed well on in-distribution data. CLIP ViT-L/14 is the best in-distribution detector, as it achieves an ID test balanced accuracy of 0.9969, an F1 score of 0.9977, and a ROC-AUC of 0.9999.
ConvNeXt-Base was second best with an ID test balanced accuracy of 0.9727.
ResNet-50 reached 0.9637, EfficientNet-B0 achieved 0.9524, and ResNet-18 was at 0.9260.
Under generator shift to SD3.5m, every model suffers a substantial performance drop on the OOD test set.
The results showed:
-
CLIP ViT-L/14: OOD balanced accuracy of 0.7829 (drop of 0.214)
-
ConvNeXt-Base: OOD balanced accuracy of 0.7643 (drop of 0.208)
-
EfficientNet-B0: OOD balanced accuracy of 0.7156 (drop of 0.237)
-
ResNet-50: OOD balanced accuracy of 0.7125 (drop of 0.251)
-
ResNet-18: OOD balanced accuracy of 0.6688 (drop of 0.257)
The paper states: "ResNet-18 experiences the largest decrease, losing 25.7 percentage points of balanced accuracy and 51.0 percentage points of AI recall. Even the strongest model, CLIP ViT-L/14, loses 21.4 percentage points of balanced accuracy."
The degradation was strongly asymmetric with false positive rates on human artwork samples remaining low while recall on SD3.5m images falls sharply.
Specifically, CLIP ViT-L/14 maintains an OOD FPR of just 0.0019, and ConvNeXt-Base maintains an OOD FPR of 0.0270.
The paper explains: the detectors remain conservative in assigning the AI label. They rarely misclassify human paintings as AI-generated, but fail to recognize a large fraction of SD3.5m images.
At the validation-selected thresholds, the five detectors miss approximately 4,300 to 5,800 of the 10,000 SD3.5m images.
The paper found that generator shift does not affect all artistic movements equally.
Ukiyo-e was consistently the easiest style for all models
with a mean OOD balanced accuracy of 0.874, while Realism is the hardest style, with no model exceeding a balanced accuracy of 0.694 for Realism.
The mean OOD AI recall for Realism was 0.337, meaning the detectors miss about two thirds of SD3.5m images in this category.
On the in-distribution test set, SD2.1 images are consistently detected at a higher rate than LDM images, with SD2.1 recall exceeding LDM recall by up to four percentage points.
The paper notes that every model achieves recall above 0.91 for both LDM and SD2.1, whereas on the OOD set, every model falls below 0.57 recall on SD3.5m images.
The qualitative analysis using Grad-CAM revealed that For ID true positives, the activation maps are usually focused on specific parts of the image rather than spread across the full canvas.
However, For OOD false negatives, the Grad-CAM maps are weaker and less clear than the maps for correct ID AI detections. The model does not focus strongly on the main subject or other important parts of the image.
The paper concludes: "These weak and peripheral attribution patterns are associated with a lower predicted AI scores. These observations are consistent with the OOD results and indicate that SD3.5m does not produce the same AI-related cues that the detector learned from LDM and SD2.1 images."
The paper concludes that strong in-distribution accuracy is not sufficient evidence of detector robustness when the underlying generator changes
and that systems that perform well on known generators may generalize poorly to newer models with different synthesis mechanisms.
The cybersecurity implications are significant: A detector that performs almost perfectly on known generators may still allow a substantial fraction of images from an emerging generator to pass as human-created.
The paper emphasizes that the risk is particularly important because it does not require a sophisticated or adaptive adversary. Simply switching to a newer generator may be sufficient to reduce detector reliability.
The authors recommend that "Image-based detection should therefore be used as one signal within a layered defensive process that also considers provenance metadata, watermark or content-credential verification, source analysis, drift monitoring, periodic evaluation against newly released generators and human review."
Future work directions include testing whether backbone fine-tuning or parameter-efficient adaptation can improve cross-generator robustness,
evaluating across additional recent image generation models,
and extending the framework from binary human-versus-AI detection to multi-generator attribution.
Improvements for AI systems
Improvements to AI Systems Based on This Paper:
- Add generator-shift robustness evaluation to detector training pipelines.
Instead of only optimizing for in-distribution accuracy, train detectors with explicit out-of-distribution (OOD) validation sets from multiple generator families (e.g., LDM, SD2.1, SD3.5m) and select model checkpoints based on OOD balanced accuracy, not just ID accuracy. This prevents overfitting to generator-specific artifacts.
- Implement uncertainty-aware detection with rejection or confidence thresholds.
Since all detectors show asymmetric failure (low false positives but high false negatives on new generators), the improved system should output a confidence score and flag low-confidence AI predictions for human review or secondary verification (e.g., watermark checks, metadata analysis) rather than making a binary decision.
- Use style-aware calibration to adjust detection thresholds per artistic style.
Given that Realism is consistently hardest (mean OOD recall 0.337) and Ukiyo-e is easiest (mean OOD balanced accuracy 0.874), the system can learn style-specific decision boundaries or apply post-hoc threshold adjustments based on the detected style of the input image, improving recall on difficult styles without increasing false positives on easy ones.
- Incorporate multi-signal fusion for layered defense.
Combine the image-based detector with non-image signals—such as provenance metadata, content-credential verification (e.g., C2PA), and generator-specific watermark detection—so that when the visual detector’s confidence is low, the system relies on auxiliary signals to make a final decision, reducing the impact of generator shift.
- Add continuous drift monitoring and periodic re-evaluation.
Deploy the detector with an automated pipeline that monitors the distribution of incoming images (e.g., feature-space drift metrics) and triggers re-training or fine-tuning when drift is detected, or when new generator versions are released. This ensures the system adapts to generator shifts over time.
- Develop a multi-generator attribution head instead of binary classification.
Extend the linear probe to output a probability distribution over known generators (LDM, SD2.1, SD3.5m, human) and use the entropy of this distribution as a signal for OOD detection. If the input is from an unseen generator, the entropy will be high, prompting a fallback to human review or additional verification.
- Use Grad-CAM-based attention regularization during training.
Train the detector to focus on the main subject of the image (e.g., by penalizing weak or peripheral activation maps on OOD examples) so that it learns more generalizable cues (e.g., texture, composition) rather than generator-specific artifacts. This can be done via auxiliary loss that encourages high-attention density on salient regions.
- Implement a two-stage cascade: coarse style classifier + specialized detectors.
First, classify the input image’s artistic style (e.g., using a lightweight style classifier). Then, route the image to a style-specific detector that has been fine-tuned or calibrated on that style’s OOD data. This mitigates the style-dependent performance gap and improves overall robustness.
What the improved AI system can do:
-
Maintain high detection accuracy on known generators while degrading gracefully (e.g., less than 10% drop) on unseen generators, instead of the 20-25% drops observed.
-
Flag low-confidence predictions for human review or metadata verification, reducing the number of missed AI images from 4,300-5,800 per 10,000 to below 1,000.
-
Detect AI-generated art in Realism style with at least 60% recall (up from 33.7%) while keeping false positive rates below 2%.
-
Automatically adapt to new generator releases via drift monitoring and periodic re-training, without requiring manual intervention.
-
Provide a confidence score and explanation (via attention maps) for each detection, enabling operators to understand why an image was classified as human or AI, and to decide when to escalate to human review.
-
Attribute an image to a specific generator family (e.g., LDM vs. SD3.5m) with high accuracy, enabling forensic analysis and better understanding of generator-specific artifacts.
Abstract
Text-to-image generative models have advanced rapidly, with modern Diffusion Transformer architectures producing images that are increasingly difficult to distinguish from human-created artwork. This development has raised significant concerns regarding copyright protection, misinformation, fraud, impersonation, and the authenticity of digital content. Most AI-art detectors are trained and evaluated on the same generator family, leaving robustness to newer architectures underexplored. In this chapter, we analyze generator shift based on a Stable Diffusion 3.5 Medium (SD3.5m) artwork dataset spanning ten art styles through reverse prompting of held-out human artwork samples. Five detectors are trained on U-Net-based latent diffusion artwork and evaluated in a zero-shot cross-generator setting on the SD3.5m dataset. Deep learning models perform strongly in-distribution but degrade under generator shift, misclassifying many SD3.5m images as human while human false positives remain low. The CLIP ViT-L/14 model performs best overall, while Grad-CAM analysis reveals weaker and more diffuse activation on false negatives. These findings highlight a generalization gap in current AI-art detectors and motivate the development of detectors as one component of a layered defense that remains reliable across rapidly evolving generative architectures.
Sources
- Demystifying MMD GANs
- RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors
- Gaussian Error Linear Units (GELUs)
- Detecting AI-generated Artwork
- The ArtBench Dataset: Benchmarking Generative Models with Artworks
- Decoupled Weight Decay Regularization
- Reverse Prompt: Cracking the Recipe Inside Text-to-Image Generation
- ArtBrain: An Explainable end-to-end Toolkit for Classification and Attribution of AI-Generated Art and Style
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models