AIGC Video Detection based on the fusion of spatial-frequency-optical flow multimodal features
cs.CV, cs.AI
Submitted: 2026-08-17
Updated: 2026-08-17
License: http://creativecommons.org/licenses/by/4.0/
The gist: The rapid evolution of generative AI (e.g., Sora, Hunyuan) makes it essential to develop effective detection strategies that can generalize across ever-evolving synthesis techniques.
Terminology
Abstract
The rapid evolution of generative AI (e.g., Sora, Hunyuan) makes it essential to develop effective detection strategies that can generalize across ever-evolving synthesis techniques. This study is motivated by the observation of a fundamental challenge in generative models: the inherent difficulty of maintaining cross-modal consistency between appearance and motion. To this end, we propose a multi-modal framework for AIGC video forgery detection tasks, named Cross-Attention based Video Forgery Detector (CrossAtt-VFD), based on joint multi-view analysis of content.Methodologically, we introduce a dual-branch architecture that simultaneously extracts spatial-frequency and optical-flow features.This approach enables the modeling of videos from complementary perceptual perspectives.The core of this process is a dedicated cross-attention mechanism, which governs the alignment of the two modalities and translates cross-modal inconsistencies into a potent diagnostic signal. This multi-modal strategy facilitates the detection of motion that is statistically inconsistent with the visual appearance of a scene. Comprehensive experimental results demonstrated that our model achieves an accuracy of 94.22%, a precision of 91.67 %,and a recall of 96.25 %, effectively verifying the advantages of the multi-modal fusion strategy.
Sources
- MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection
- ViTransPAD: Video Transformer using convolution and self-attention for Face Presentation Attack Detection
- VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models