DDL: A Large-Scale Dataset for Deepfake Detection and Localization in Diversified Real-World Scenarios
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "DDL: A Large-Scale Dataset for Deepfake Detection and Localization in Diversified Real-World Scenarios".
Tom: The DDL dataset introduces a large-scale, diverse,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the creators for a minute. The paper is called "DDL: A Large-Scale Dataset for Deepfake Detection and Localization in Diversified Real-World Scenarios," written by Changtao Miao, Yi Zhang, Weize Gao, Zhiya Tan, Weiwei Feng, Man Luo, Jianshu Li, Ajian Liu, Yunfeng Diao.
Jane: Those are a lot of names for one paper. What does this title actually tell us about the scope of the work?
Tom: It tells us it’s not just about one kind of fake; it covers a huge variety of techniques—about eighty different deepfake methods, including everything from old GANs to newer diffusion models and commercial software.
Lu: That breadth is really impressive. They included generation architectures like VAEs, NeRFs, and even autoregressive models alongside popular ones like StyleGAN and Kling two point one <ref:2506.23292#pg1>.
Meng: So it’s not just a collection of images; it’s a comprehensive benchmark covering diverse technologies to stress-test detection systems across the board.
Lalam: And they cover different modes of manipulation too, like face swapping, reenactment, and fullface synthesis in the spatial domain, plus deletion or replacement in the temporal domain.
Tom: It’s really about making sure that whatever detection tool we build can handle a wide spectrum of forgery styles and techniques.
Jane: That diversity is key because it shows that a model trained on one type of fake might completely fail when faced with another, which is what we want to see in testing.
The paper's summary: Tom: Now let’s look at what DDL actually delivers according to the authors. They summarize the dataset as being over one point four million forged samples, covering a vast range of scenarios and modalities including image, audio, and video.
Jane: So when they talk about those one point four million samples, that’s not just a big number; it means there’s enough variety to really train something robust without running out of challenging examples too quickly.
Lu: They emphasize that the main reason existing datasets fall short is that they provide only image-level or video-level binary labels, and DDL specifically addresses this by adding fine-grained annotations.
Meng: They are providing precise spatial masks for where the forgery is located and temporal segments for when it happens, which directly supports those localization tasks we talked about.
Lalam: They even provide one point one eight million spatial masks and zero point two three million temporal segments, which is a huge amount of detail to work with when training segmentation or tracking models.
Tom: That’s the meat of it—they are shifting the focus from just spotting a fake to precisely mapping out the manipulation itself across different media types.
The paper's improvements: Tom: So, what are the actual improvements they suggest for using this dataset? They aren't just presenting data; they’re showing how this new annotation level helps researchers do things that were previously impossible or very hard.
Jane: The main improvement is enabling spatial and temporal forgery localization tasks. This means AI systems can now output precise one point one eight million spatial masks and zero point two three million temporal segments for various scenarios like single-face, multi-face, and audio visual content.
Lu: That capability directly enhances interpretability; we get to see exactly "where" the manipulation is occurring in the sample, which is much more informative than a simple detection score.
Meng: And they’ve also put in a lot of work on data quality through human-in-the-loop oversight, where experts review prompts and then screen samples themselves using criteria like realism artifacts and temporal coherence.
Lalam: That human oversight step is really important because it ensures the annotations are accurate, which is crucial when you’re training high-stakes models on this kind of detailed labeling.
Tom: They also designed a test set that includes twenty-seven types of real-world perturbations, like color corruption and weather changes for images, or H <ref:2506.23292#pg2>.two hundred sixty-four compression and Gaussian noise for audio visual content.
Conclusion: Jane: So to wrap up on the DDL paper, the big picture is that they’ve created a large-scale dataset that moves deepfake research beyond simple detection into precise localization across multiple modalities.
Tom: They achieved this by including over eighty deepfake methods and providing very detailed spatial and temporal annotations on over one point four million samples, which gives researchers the tools to build much better detection models.
Lu: The implication is that next-generation deepfake detection won't just be about classification anymore; it will be about understanding the mechanics of the forgery itself at a granular level.
Meng: From an engineering standpoint, having this diverse and well-annotated data helps us build models that are less likely to overfit to specific generation styles because they are tested against a much wider distribution.
Lalam: And with the human oversight included, they’ve also built a quality pipeline for sure, which is something engineers always appreciate when dealing with massive datasets like this.
Jane: It’s a really solid contribution to the field, giving us the necessary tools to analyze and fight sophisticated manipulations in real-world scenarios.
AntGroup Institute of Automation, Chinese Academy of Sciences, Hefei University of Technology
cs.CV
Submitted: 2025-06-29
Updated: 2026-10-08
Project page: https://deepfake-workshop-ijcai2025.github.io/main/index.html
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 89/100
The gist: The DDL dataset introduces a large-scale, diverse, multi-modal deepfake detection and localization benchmark to address the limitations of existing datasets by providing fine-grained spatial and
Key concepts
- Comprehensive Deepfake Methods
- The dataset includes a wide variety of deepfake creation techniques, spanning 7 generation architectures like GANs and Diffusion models. It covers 80 distinct methods, ensuring the benchmark tests detection against a broad spectrum of forgery technologies, from common to emerging models.
- Fine-grained Forgery Annotations
- The DDL provides extremely detailed labels for forged content. This includes precise spatial masks for where manipulation occurred in images and temporal segment labels for when manipulations happened in videos, significantly boosting the accuracy of localization research.
- Multi-modal Content (DDL-AV)
- The dataset supports both unimodal image data (DDL-I) and multi-modal audio-visual content (DDL-AV). DDL-I is used for spatial forgery localization, while DDL-AV is specifically designed for temporal forgery localization, making the benchmark versatile.
- Human-in-the-Loop Oversight
- To ensure high quality and accuracy, human experts reviewed the prompts used to generate samples and then manually screened and annotated the resulting deepfakes. This human judgment compensates for automated system errors, ensuring reliable training data.
Terminology
Summary
The DDL dataset introduces a large-scale, diverse, multi-modal deepfake detection and localization benchmark to address the limitations of existing datasets by providing fine-grained spatial and temporal annotations for complex real-world forgery scenarios. The gist: This paper proposes a novel large-scale deepfake detection and localization (DDL) dataset containing over 1.4M+ forged samples encompassing up to 80 distinct deepfake methods, which provides crucial support for building next-generation deepfake detection, localization, and interpretability methods.
Dataset Construction and Innovations
The DDL dataset was constructed to overcome the limitations of current datasets which often lack fine-grained annotations of manipulation The DDL design incorporates four key innovations including (1) Comprehensive Deepfake Methods covering 7 different generation architectures and a total of 80 methods, (2) Varied Manipulation Modes incorporating 7 classic and 3 novel forgery modes, (3) Diverse Forgery Scenarios and Modalities including 3 scenarios and 3 modalities, and (4) Fine-grained Forgery Annotations providing precise spatial masks or temporal segments Specifically, the dataset encompasses both unimodal images (DDL-I) and multi-modal audio-visual (DDL-AV) content specifically designed for spatial forgery localization and temporal forgery localization tasks respectively The total number of forged samples in the DDL dataset is over 1.4M+
Comprehensive Deepfake Methods and Modalities
The dataset ensures technological diversity by including 80 state-of-the-art Deepfake techniques spanning from common GANs [6] and Diffusion models [8] to emerging architectures such as VAEs [12], Normalizing Flows [13], NeRFs [20], Autoregressive [25] models, and popular commercial software Furthermore, it encompasses both visual and audio modalities enriching the dataset’s technological diversity In terms of manipulation modes, DDL covers face swapping, face reenactment, fullface synthesis, and face editing in the spatial domain In the temporal domain it includes deletion, replacement, and insertion operations of forged content Notably it introduces hybrid face forgery audio-visual asynchronous manipulation and audio-visual full synthesis modes for the first time increasing complexity and realism
Diverse Scenarios and Fine-Grained Annotations
DDL covers single-face, multi-face, and audio-visual scenarios while incorporating audio, image, and video modalities simulating complex real-world forgery content The dataset includes 80 Deepfake techniques resulting in a significantly more diverse and challenging benchmark For fine-grained annotations we provide spatial forgery region masks and temporal forgery segment labels including precise 1.18M+ spatial masks and 0.23M+ temporal segments These detailed annotations significantly enhance the research capabilities for forgery localization tasks
Data Quality through Human-in-the-Loop Oversight
To ensure deepfake sample quality compliance, and annotation accuracy we incorporate human experts’ intelligence and judgment to compensate for the shortcomings of automated systems This oversight involves two main steps (3.1.1) Generation Prompts Curation where human experts review LLM-generated prompts checking their accuracy, completeness, clarity, and alignment with predefined generation specifications and ethical standards (3.1.2) Deepfake Sample Quality and Annotation where human experts conduct quality screening and annotation of generated samples to ensure that datasets used for training and testing are high-quality, accurate, and representative Experts assess deepfake samples using both subjective and objective criteria including realism artifacts naturalness consistency of expressions movements, and temporal coherence continuity within videos
Real-World Perturbations and Out-of-Distribution Testing
To simulate real-world transmission we design and apply 27 types of perturbation methods to the test set samples For the image modality perturbations are categorized into three groups color corruption, and weather with each category comprising 8 distinct perturbation methods For the audio-visual modality we employ a joint perturbation strategy including H.264-based compression, Gaussian noise, and reverberation blur The DDL dataset deliberately isolates the distributions of the training and test sets from two perspectives sources of real data and generation model types used enabling the construction of out-of-distribution test sets As shown in Figure 6(b) the source overlap rate of test samples in DDL is only 21.05% compared to 98.98% for DF40, and a full 100% overlap for both AV-DF1M and ForgeryNet
Conclusion
The main contributions are three-folds We propose a large-scale diverse and multi-modal dataset DDL which contains audio-visual content and finegrained forgery annotations with 1.18M+ spatial masks and 0.
Improvements for AI systems
-
A large-scale, diverse, multi-modal dataset (DDL) will enable deepfake detection models to move beyond simple binary classification by supporting
spatial forgery localization and temporal forgery localization tasks.
This allows AI systems to provideprecise 1.18M+ spatial masks and 0.23M+ temporal segments,
significantly enhancing interpretability by showing exactlywhere
andwhen
manipulation occurred in a forged sample. -
The unified deepfake generation pipeline, driven by LLMs and humans, will allow AI systems to generate highly complex, realistic forgeries across varied modalities (audio, image, video). This capability enables the creation of new forgery types such as
hybrid face forgery (HFF) mode
andaudio-visual full synthesis (AVFS) mode,
providing a more challenging benchmark for detection models. -
AI systems will be equipped to perform real-world perturbation analysis by applying
27 types of perturbation methods
to test set samples, including color, corruption, and weather changes. This allows the system to assess the robustness of deepfake detection models against common transmission artifacts likeH.264-based compression, Gaussian noise, and reverberation blur.
-
Detection and localization systems will gain superior generalization by being trained on data with isolated distributions between training and testing sets. By isolating distributions from
the sources of real data and the generation model types used,
the system can be tested onout-of-distribution test sets,
preventing overfitting to specific forgery styles present in the training set.
Abstract
Recent advances in AIGC have exacerbated the misuse of malicious deepfake content, making the development of reliable deepfake detection methods an essential means to address this challenge. Although existing deepfake detection models demonstrate outstanding performance in detection metrics, most methods only provide simple binary classification results, lacking interpretability. Recent studies have attempted to enhance the interpretability of classification results by providing spatial manipulation masks or temporal forgery segments. However, due to the limitations of forgery datasets, the practical effectiveness of these methods remains suboptimal. The primary reason lies in the fact that most existing deepfake datasets contain only binary labels, with limited variety in forgery scenarios, insufficient diversity in deepfake types, and relatively small data scales, making them inadequate for complex real-world scenarios. To address this predicament, we construct a novel large-scale deepfake detection and localization (DDL) dataset containing 1.4M+ forged samples and encompassing 80 distinct deepfake methods. The DDL design incorporates four key innovations: (1) Comprehensive Deepfake Methods (covering 7 different generation architectures and a total of 80 methods), (2) Varied Manipulation Modes (incorporating 7 classic and 3 novel forgery modes), (3) Diverse Forgery Scenarios and Modalities (including 3 scenarios and 3 modalities), and (4) Fine-grained Forgery Annotations (providing 1.18M+ precise spatial masks and 0.23M+ precise temporal segments). Through these improvements, our DDL not only provides a more challenging benchmark for complex real-world forgeries but also offers crucial support for building next-generation deepfake detection, localization, and interpretability methods.
Sources
- Diffusion Deepfake
- The DeepFake Detection Challenge (DFDC) Dataset
- DeeperForensics Challenge 2020 on Real-World Face Forgery Detection: Methods and Results
- FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset
- Multi-spectral Class Center Network for Face Manipulation Detection and Localization
- Mixture-of-Noises Enhanced Forgery-Aware Predictor for Multi-Face Manipulation Detection and Localization
- Robustness and Generalizability of Deepfake Detection: A Study with Diffusion Models
- DiffusionFake: Enhancing Generalization in Deepfake Detection via Guided Stable Diffusion
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models