DDL: A Large-Scale Dataset for Deepfake Detection and Localization in Diversified Real-World Scenarios
summary
The gist
The DDL dataset introduces a large-scale, diverse, multi-modal deepfake detection and localization benchmark to address the limitations of existing datasets by providing fine-grained spatial and
In short
The DDL dataset was created to improve deepfake detection and localization by addressing limitations in existing benchmarks. It provides over 1.4 million forged samples covering 80 different deepfake methods across various visual and audio modalities, complete with fine-grained spatial and temporal annotations, allowing researchers to build more robust detection tools.
Key concepts
- Comprehensive Deepfake Methods
- The dataset includes a wide variety of deepfake creation techniques, spanning 7 generation architectures like GANs and Diffusion models. It covers 80 distinct methods, ensuring the benchmark tests detection against a broad spectrum of forgery technologies, from common to emerging models.
- Fine-grained Forgery Annotations
- The DDL provides extremely detailed labels for forged content. This includes precise spatial masks for where manipulation occurred in images and temporal segment labels for when manipulations happened in videos, significantly boosting the accuracy of localization research.
- Multi-modal Content (DDL-AV)
- The dataset supports both unimodal image data (DDL-I) and multi-modal audio-visual content (DDL-AV). DDL-I is used for spatial forgery localization, while DDL-AV is specifically designed for temporal forgery localization, making the benchmark versatile.
- Human-in-the-Loop Oversight
- To ensure high quality and accuracy, human experts reviewed the prompts used to generate samples and then manually screened and annotated the resulting deepfakes. This human judgment compensates for automated system errors, ensuring reliable training data.
Terminology used across episodes
This episode discusses
- DDL: A Large-Scale Dataset for Deepfake Detection and Localization in Diversified Real-World Scenarios · Paper Radio
- Diffusion Deepfake
- The DeepFake Detection Challenge (DFDC) Dataset
- DeeperForensics Challenge 2020 on Real-World Face Forgery Detection: Methods and Results
- FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset
- Multi-spectral Class Center Network for Face Manipulation Detection and Localization
- Mixture-of-Noises Enhanced Forgery-Aware Predictor for Multi-Face Manipulation Detection and Localization
- Robustness and Generalizability of Deepfake Detection: A Study with Diffusion Models
- DiffusionFake: Enhancing Generalization in Deepfake Detection via Guided Stable Diffusion
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
The paper
DDL: A Large-Scale Dataset for Deepfake Detection and Localization in Diversified Real-World Scenarios · Read on arXiv
AntGroup Institute of Automation, Chinese Academy of Sciences, Hefei University of Technology
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "DDL: A Large-Scale Dataset for Deepfake Detection and Localization in Diversified Real-World Scenarios".
Tom: The DDL dataset introduces a large-scale, diverse,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the creators for a minute. The paper is called "DDL: A Large-Scale Dataset for Deepfake Detection and Localization in Diversified Real-World Scenarios," written by Changtao Miao, Yi Zhang, Weize Gao, Zhiya Tan, Weiwei Feng, Man Luo, Jianshu Li, Ajian Liu, Yunfeng Diao.
Jane: Those are a lot of names for one paper. What does this title actually tell us about the scope of the work?
Tom: It tells us it’s not just about one kind of fake; it covers a huge variety of techniques—about eighty different deepfake methods, including everything from old GANs to newer diffusion models and commercial software.
Lu: That breadth is really impressive. They included generation architectures like VAEs, NeRFs, and even autoregressive models alongside popular ones like StyleGAN and Kling two point one <ref:2506.23292#pg1>.
Meng: So it’s not just a collection of images; it’s a comprehensive benchmark covering diverse technologies to stress-test detection systems across the board.
Lalam: And they cover different modes of manipulation too, like face swapping, reenactment, and fullface synthesis in the spatial domain, plus deletion or replacement in the temporal domain.
Tom: It’s really about making sure that whatever detection tool we build can handle a wide spectrum of forgery styles and techniques.
Jane: That diversity is key because it shows that a model trained on one type of fake might completely fail when faced with another, which is what we want to see in testing.
The paper's summary: Tom: Now let’s look at what DDL actually delivers according to the authors. They summarize the dataset as being over one point four million forged samples, covering a vast range of scenarios and modalities including image, audio, and video.
Jane: So when they talk about those one point four million samples, that’s not just a big number; it means there’s enough variety to really train something robust without running out of challenging examples too quickly.
Lu: They emphasize that the main reason existing datasets fall short is that they provide only image-level or video-level binary labels, and DDL specifically addresses this by adding fine-grained annotations.
Meng: They are providing precise spatial masks for where the forgery is located and temporal segments for when it happens, which directly supports those localization tasks we talked about.
Lalam: They even provide one point one eight million spatial masks and zero point two three million temporal segments, which is a huge amount of detail to work with when training segmentation or tracking models.
Tom: That’s the meat of it—they are shifting the focus from just spotting a fake to precisely mapping out the manipulation itself across different media types.
The paper's improvements: Tom: So, what are the actual improvements they suggest for using this dataset? They aren't just presenting data; they’re showing how this new annotation level helps researchers do things that were previously impossible or very hard.
Jane: The main improvement is enabling spatial and temporal forgery localization tasks. This means AI systems can now output precise one point one eight million spatial masks and zero point two three million temporal segments for various scenarios like single-face, multi-face, and audio visual content.
Lu: That capability directly enhances interpretability; we get to see exactly "where" the manipulation is occurring in the sample, which is much more informative than a simple detection score.
Meng: And they’ve also put in a lot of work on data quality through human-in-the-loop oversight, where experts review prompts and then screen samples themselves using criteria like realism artifacts and temporal coherence.
Lalam: That human oversight step is really important because it ensures the annotations are accurate, which is crucial when you’re training high-stakes models on this kind of detailed labeling.
Tom: They also designed a test set that includes twenty-seven types of real-world perturbations, like color corruption and weather changes for images, or H <ref:2506.23292#pg2>.two hundred sixty-four compression and Gaussian noise for audio visual content.
Conclusion: Jane: So to wrap up on the DDL paper, the big picture is that they’ve created a large-scale dataset that moves deepfake research beyond simple detection into precise localization across multiple modalities.
Tom: They achieved this by including over eighty deepfake methods and providing very detailed spatial and temporal annotations on over one point four million samples, which gives researchers the tools to build much better detection models.
Lu: The implication is that next-generation deepfake detection won't just be about classification anymore; it will be about understanding the mechanics of the forgery itself at a granular level.
Meng: From an engineering standpoint, having this diverse and well-annotated data helps us build models that are less likely to overfit to specific generation styles because they are tested against a much wider distribution.
Lalam: And with the human oversight included, they’ve also built a quality pipeline for sure, which is something engineers always appreciate when dealing with massive datasets like this.
Jane: It’s a really solid contribution to the field, giving us the necessary tools to analyze and fight sophisticated manipulations in real-world scenarios.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck