Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation
Yichun Yeh, Yiheng Li, Xiaobo Hu, Zhen Lei, Yang Yang
cs.CV, cs.AI
Submitted: 2026-08-08
Updated: 2026-08-11
Comments: accepted by ACM MM 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Fake news increasingly relies on cross-modal image-text forgeries, making transparent and verifiable reasoning chains an urgent need for Detecting and Grounding Multi-Modal Media Manipulation (DGM4).
Terminology
Abstract
Fake news increasingly relies on cross-modal image-text forgeries, making transparent and verifiable reasoning chains an urgent need for Detecting and Grounding Multi-Modal Media Manipulation (DGM4). Existing methods produce black-box detection results without any decision rationale, limiting their reliability in forensic practice. Multi-modal Large Language Models (MLLMs) offer a natural path toward explainability, but applying them to DGM4 raises two difficulties. First, models tend to generate explanations disconnected from predicted evidence locations, producing unverified attribution. Second, enforcing evidence-conclusion consistency requires active optimization, yet uniform training signals fail to distinguish localization tokens from classification tokens, making multi-head joint training unreliable. We propose a multi-modal manipulation detector based on an Evidence-Grounded Forensic Reasoning (EFR) framework. EFR introduces an Anchor-and-Verify reasoning chain that enforces modality-isolated perception before cross-modal comparison, with conclusion coordinates as explicit anchors to which downstream evidence must spatially correspond. A verifiable reward system then enforces evidence-conclusion consistency during training, while a Modality-Decoupled Advantage (MDA) routing mechanism mitigats credit misassignment across prediction tasks. Experiments show that EFR achieves state-of-the-art performance while producing structured forensic reasoning records that explicitly bind explanations to evidence.
Sources
- Qwen3-VL Technical Report
- LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation
- MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language Models
- FORGE: Forensic Reasoning with Grounded Evidence
- Seeing Before Reasoning: A Unified Framework for Generalizable and Explainable Fake Image Detection
- ForgeryGPT: A Multimodal LLM for Interpretable Image Forgery Detection and Localization
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
- Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis
- FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Faces
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-Reflection
- DGM4+: Dataset Extension for Global Scene Inconsistency
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- MARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models
- FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models