Synthetic Image Detection with CLIP: Understanding and Assessing Predictive Cues
cs.CV
Submitted: 2026-02-12
Updated: 2026-08-15
Code: https://github.com/marco-willi/clip-cues
Terminology
Sources
- GenAI Mirage: The Impostor Bias and the Deepfake Detection Challenge in the Era of Artificial Illusions
- Reproducible scaling laws for contrastive language-image learning
- Are CLIP features all you need for Universal Synthetic Image Origin Attribution?
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- ForensicTransfer: Weakly-supervised Domain Adaptation for Forgery Detection
- Raising the Bar of AI-generated Image Detection with CLIP
- Density estimation using Real NVP
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- Improving Interpretability and Robustness for the Detection of AI-Generated Images
- Interpreting CLIP's Image Representation via Text-Based Decomposition
- Denoising Diffusion Probabilistic Models
- ASAP: Interpretable Analysis and Summarization of AI-generated Image Patterns at Scale
- Imagen 3
- CLIPping the Deception: Adapting Vision-Language Models for Universal Deepfake Detection
- Auto-Encoding Variational Bayes
- Leveraging Representations from Intermediate Encoder-blocks for Synthetic Image Detection
- Flow Matching for Generative Modeling
- Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated Images
- Towards Universal Fake Image Detectors that Generalize Across Generative Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models