ForensicZoom: Adaptive Visual Inspection with Multimodal LLMs for Industrial-Grade Face Forgery Detection
cs.CV
Submitted: 2026-09-15
Updated: 2026-09-15
Terminology
Sources
- Qwen3-VL Technical Report
- Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024
- Combining EfficientNet and Vision Transformers for Video Deepfake Detection
- Generative Adversarial Networks
- Denoising Diffusion Probabilistic Models
- FFAA: Multimodal Large Language Model based Explainable Open-World Face Forgery Analysis Assistant
- Machine Learning-based Approach for Ex-post Assessment of Community Risk and Resilience Based on Coupled Human-infrastructure Systems Performance
- Explainable Deepfake Detection with RL Enhanced Self-Blended Images
- A Style-Based Generator Architecture for Generative Adversarial Networks
- ForgeryGPT: A Multimodal LLM for Interpretable Image Forgery Detection and Localization
- Detecting GAN-generated Imagery using Color Cues
- Towards Universal Fake Image Detectors that Generalize Across Generative Models
- Proximal Policy Optimization Algorithms
- How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
- Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
- CNN-generated images are surprisingly easy to spot... for now
- ForgeryVCR: Visual-Centric Reasoning via Efficient Forensic Tools in MLLMs for Image Forgery Detection and Localization
- V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
- Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
- MARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models