FORGE: Forensic Reasoning with Grounded Evidence
cs.CV, cs.AI
Submitted: 2025-03-20
Updated: 2026-09-16
Comments: Accepted at EMNLP 2026 Findings
License: http://creativecommons.org/licenses/by/4.0/
The gist: Forensic deepfake analysis demands more than binary classification: investigators need region-grounded natural language explanations they can verify against the image.
Terminology
Abstract
Forensic deepfake analysis demands more than binary classification: investigators need region-grounded natural language explanations they can verify against the image. Multimodal large language models (MLLMs) are a natural fit, but pretrained MLLMs fail systematically, producing globally coherent text that misses the small localized cues defining manipulations. We argue this is an inductive bias problem rather than a capacity issue: the image-text contrastive objective training MLLM visual encoders optimizes for whole-image semantic summaries, not patch-level forensic detail. The same mismatch explains why prior deepfake reasoning methods target either face manipulation or fully AI-generated content, never both. We propose FORGE, which addresses the mismatch by routing a second visual stream into the language model from a Vision-Only Model (VOM) trained on dense patch prediction rather than image-text alignment. The MLLM's native encoder and the VOM operate on a shared patch grid, which lets us interleave their tokens with preserved spatial correspondence; we show this beats naive concatenation. A two-stage adapter training protocol (generic image-caption alignment, then joint task-specific optimization) prevents the localized stream from overfitting to training-domain manipulations. Across face-manipulated and fully synthetic content, FORGE produces region-referential explanations answering fine-grained attribute queries ("Does the eyes/nose/mouth look real or fake?") and substantially outperforms in-domain baselines on cross-domain evaluations; region-specific evaluation and human studies confirm explanation faithfulness.
Sources
- GPT-4 Technical Report
- Gemini: A Family of Highly Capable Multimodal Models
- PaliGemma: A versatile 3B VLM for transfer
- DeMamba: AI-Generated Video Detection on Million-Scale GenVideo Benchmark
- X2-DFD: A framework for eXplainable and eXtendable Deepfake Detection
- Can We Leave Deepfake Data Behind in Training Deepfake Detector?
- Vision-Language Models Can't See the Obvious
- Cross-Modal Adapter for Vision-Language Retrieval
- FaceShifter: Towards High Fidelity And Occlusion Aware Face Swapping
- Exposing DeepFake Videos By Detecting Face Warping Artifacts
- DINOv2: Learning Robust Visual Features without Supervision
- SHIELD : An Evaluation Benchmark for Face Spoofing and Forgery Detection with Multimodal Large Language Models
- PaliGemma 2: A Family of Versatile VLMs for Transfer
- ModelScope Text-to-Video Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models