Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection".
Elias: The gist: The proposed Detection-Guided Adaptive Purification (DGAP) framework is a diffusion-based defense that adjusts purification strength per input based on its score shift relative to the detector's response,
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: We've just talked about how this paper, "Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection," tackles the vulnerability of existing deepfake detectors to adversarial audio perturbations <ref:2610.10752#pg1>.
Elias: The thesis is that instead of using a single purification strength for all inputs, you can adjust that strength per input based on how much the detector's score shifts when you apply a light purification step <ref:2610.10752#pg2>.
Priya: So, what's the main claim they are making about this adaptive process? Is it just that it works better in testing, or is there something deeper there?
Nadia: The paper claims this mechanism leaves benign inputs nearly unaffected while concentrating purification effort where it is most needed to neutralize adversarial perturbations <ref:2610.10752#pg2>.
Elias: They are building on the observation that a light purification step causes a significant score shift for adversarial inputs compared to benign ones <ref:2610.10752#pg1>.
Priya: That score shift is used as a reference-free indicator, meaning the system doesn't need any clean reference audio to know if it needs to do more purification <ref:2610.10752#pg2>.
Nadia: Exactly. They introduce a gating mechanism that flags an input if its score gap exceeds a threshold set on benign inputs, and only then does it apply the stronger purification <ref:2610.10752#pg3>.
Elias: The main contribution is this Adaptive Purification Framework, which purifies each input with a diffusion model at a strength selected from the detector’s own response without modifying or retraining the detector <ref:2610.10752#pg2>.
Priya: So, what does this mean for us in terms of real-world application? Is it just academic stuff about improving test scores?
Nadia: It matters because existing methods often require retraining the detector or introduce extra distortion, but DGAP modifies the input before classification to improve robustness <ref:2610.10752#pg3>.
Elias: They are also looking at related ideas that have been explored in speech detection, like denoising or self-supervised resynthesis, but this is a specific application of diffusion models to this audio problem <ref:2610.10752#pg3>.
Priya: I’m interested in the idea that it operates on input modification rather than retraining the detector itself, which seems like a cleaner way to defend against attacks <ref:2610.10752#pg3>.
Nadia: That's the core of its appeal; modifying the input before classification allows you to maintain a pretrained detector while boosting its defense capabilities <ref:2610.10752#pg2>.
Elias: It's a way to make the existing detection system more resilient by intelligently choosing the level of signal cleanup for each specific audio sample <ref:2610.10752#pg3>.
Priya: So, to sum up, it’s about creating a dynamic defense that treats different inputs differently based on their potential threat level <ref:2610.10752#pg3>.
Conclusion: Nadia: So, wrapping up this discussion on "Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection," we’ve seen how they propose a method that adapts purification strength per input <ref:2610.10752#pg1>.
Elias: The authors are Muhammed Salih Kayhan and Qiben Yan from Michigan State University, and the implications hinge on their success in achieving the strongest overall defense performance across all detectors <ref:2610.10752#pg3>.
Priya: It seems like this research points toward a future where audio security isn't just about one static defense, but having a system that reacts intelligently to the specific characteristics of each incoming sample <ref:2610.10752#pg3>.
Nadia: That's right. The key is using the detector's own response to guide the purification strength dynamically, which seems like a much more intelligent way to handle adversarial inputs than fixed transformations <ref:2610.10752#pg3>.
Elias: It shows that diffusion-based purification can be used not just for general noise reduction, but as a targeted defense against targeted manipulation in audio systems <ref:2610.10752#pg3>.
Priya: So, for the listener who's just hearing this, it means that when you use an AI system to check audio authenticity, this paper suggests a more sophisticated layer of defense is possible <ref:2610.10752#pg3>.
Nadia: It implies that the next step in robust audio detection might involve integrating these kinds of adaptive input processing techniques <ref:2610.10752#pg3>.
Muhammed Salih Kayhan, Qiben Yan
Michigan State University
cs.CR, cs.SD
Submitted: 2026-10-07
Updated: 2026-10-07
Comments: Accepted at the 4th EAI International Conference on Security and Privacy in Cyber-Physical Systems and Smart Vehicles (EAI SmartSP 2026)
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
The gist: The gist: The proposed Detection-Guided Adaptive Purification (DGAP) framework is a diffusion-based defense that adjusts purification strength per input based on its score shift relative to the
Key concepts
- Probe Stage
- This initial stage applies a light diffusion process to an incoming audio sample. By comparing the detector scores of the original audio and this lightly purified version, the framework creates a reference-free indicator that shows how much an input is likely adversarial, as benign inputs are perturbed less significantly.
- Gate Stage
- This stage acts as a filter based on the score gap measured in the probe. If an input's score gap exceeds a threshold set by benign data, it is flagged as adversarial and proceeds to purification. Inputs below this threshold are passed through unchanged to preserve their original features.
- Purify Stage
- Only inputs flagged by the gate stage undergo a substantially stronger diffusion purification process. This targeted, high-strength purification is applied specifically to potential deepfakes, ensuring that the final audio fed to the detector is robust against adversarial manipulations.
Terminology
Summary
The gist: The proposed Detection-Guided Adaptive Purification (DGAP) framework is a diffusion-based defense that adjusts purification strength per input based on its score shift relative to the detector's response, achieving the best overall defense performance across all detectors while leaving benign inputs nearly unaffected
How it works
The DGAP framework is a diffusion-based defense that adjusts purification strength per input based on its score shift relative to the detector's response, achieving the best overall defense performance across all detectors while leaving benign inputs nearly unaffected The framework first applies a light purification step and observes how the detector’s response changes If the score shifts significantly, the input is flagged and subjected to stronger purification, while the remaining inputs pass unchanged This mechanism uses a reference-free indicator of adversarial manipulation
by comparing the detector scores of the original audio and its lightly purified version
Key components of DGAP
The framework operates through three stages: probe, gate, and purify.
-
Probe Stage: DGAP applies a brief forward-reverse diffusion process at a low noise level to an incoming utterance and measures the resulting shift between the detector scores of the original waveform and the purified copy This score gap serves as a
per-input, reference-free indicator of adversarial manipulation
because a light purification perturbs an adversarial input far more than that of a benign one -
Gate Stage: An input whose gap exceeds a threshold calibrated on benign inputs is treated as adversarial, whereas an input below the threshold is forwarded unchanged to preserve the discriminative cues the detector exploits The gate requires no clean reference and no knowledge of the attack
-
Purify Stage: Only a flagged input is purified at a substantially stronger level, and the resulting waveform is passed to the detector for the final decision
Evaluation and Results
The framework was evaluated against three white-box attack settings across three deepfake detectors, comparing it with nine existing defenses. The results show that DGAP achieves the strongest overall defense performance across all detectors while leaving benign inputs nearly unaffected
Table II reports that the undefended detectors are highly vulnerable, with EERadv reaching 61.92% on AASIST and 57.49% on Res-TSSDNet The transformation-based defenses recover only partially and inconsistently across attacks
Analysis of the Gate Mechanism
The gate statistics explain the gap to the non-adaptive setting, showing that under a defense-aware attack, only 33.3% and 34.2% of adversarial utterances are flagged, against 5.8% and 4.2% of benign inputs The adversary concentrates the score gap of adversarial inputs just below the threshold η, so that the probe barely shifts the detector The gate evasion is noted as the principal failure mode of our method under adaptive attack
Practical Considerations
The running time analysis shows that every input incurs the light probe purification and scoring, whereas the strong purification is incurred only when the gate flags the input The per-utterance time is therefore the probe cost plus the strong-purification cost weighted by the flag rate This structure ensures that benign inputs pass through almost untouched and EERclean stays within less than one percent The main limitation is that under a defense-aware adaptive attack, the main failure mode is gate evasion, where the adversary can optimize the perturbation so that the score gap remains below the gating threshold and the strong purification is bypassed Future work should investigate stronger gating strategies, such as randomized probe levels or multiple probe strengths The paper concludes that DGAP achieves the strongest overall defense performance compared to existing defenses while remaining effective under a defense-aware adaptive attack
ACKNOWLEDGMENT
This work was supported in part by the U.S. National Science Foundation grant CNS-2310207
REFERENCES
[1] M. U. Farooq, A. Khan, I. U. Haq, and K. M. Malik, “Securing social media against deepfakes using identity, behavioral, and geometric signatures,” arXiv preprint arXiv:2412.05487, 2024
[3] G. Hua, A. B. J. Teoh, and H. Zhang, “Towards end-to-end synthetic speech detection,” IEEE Signal Processing Letters, vol. 28, pp. 1265–1269, 2021
[4] H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, “End-to-end anti-spoofing with RawNet2,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021
[8] C. Guo, M. Rana, M. Cissé, and L. van der Maaten, “Countering adversarial images using input transformations,” in International Conference on Learning Representations, 2018
[9] C. Xie, J. Wang, Z. Zhang, Z. Ren, and A. Yuille, “Mitigating adversarial effects through randomization,” in International Conference on Learning Representations, 2018
[10] T. Bai, J. Luo, J. Zhao, B. Wen, and Q. Wang, “Recent advances in adversarial training for adversarial robustness,” in Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, 2021
[13] S. Wu, J. Wang, W. Ping, W. Nie, and C. Xiao, “Defending against adversarial audio via diffusion model,” in International Conference on Learning Representations, 2023
[14] H. Guo, G. Wang, B. Chen, Y. Wang, X. Zhang, X. Chen, Q. Yan, and L. Xiao, “WavePurifier: Purifying audio adversarial examples via hierarchical diffusion models,” in Proceedings of the 30th Annual International Conference on Mobile Computing and Networking, 2024
[18] X. Cheng, M. Xu, and T. F. Zheng, “Replay detection using CQT-based modified group delay feature and ResNeWt network in ASVspoof 2019,” in 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC). IEEE, 2019
[36] A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” in International Conference on Learning Representations, 2018
[37] X. Cheng, K. Fu, and F. Farnia, “Stability and generalization in free adversarial training,” Transactions on Machine Learning Research, 2024
[38] J. H. Metzen, T. Genewein, V. Fischer, and B. Bischoff, “On detecting adversarial perturbations,” in International Conference on Learning Representations, 2017
[39] K. Lee, K. Lee, H. Lee, and J. Shin, “A simple unified framework for detecting out-of-distribution samples and adversarial attacks,” in Advances in Neural Information Processing Systems, vol.
Improvements for AI systems
-
Adaptive Purification Framework (DGAP): The system can adjust purification strength per input by using
the resulting score shift as a reference-free indicator of adversarial manipulation,
allowing it to pass benign inputsunchanged
while applyingstronger purification before final detection
only to flagged inputs, thus achieving the objective:(1) restoring detection on adversarial spoofing audios; (2) preserving spoof artifacts, so that a clean (non-adversarial) spoof is not turned into a bonafidelooking sample; and (3) preserving bonafide speech.
-
Gating Mechanism: The system introduces
a reference-free gating mechanism that flags adversarially perturbed inputs before final detection
by comparing the detector scores of the original audio and its lightly purified version, specifically flagging inputs wherethe gap exceeds a threshold calibrated on benign inputs.
-
Robustness against Fixed Detectors: The framework is evaluated against
three deepfake detectors
and achievesthe strongest overall defense performance across all detectors while leaving benign inputs nearly unaffected,
demonstrating effectiveness even under thedefense-aware adaptive attack.
-
Parameter Optimization for Generalization: The system selects a single attack-agnostic configuration by minimizing the development objective, which indicates that
retuning per attack would instead require knowledge of the attack type, norm, and perturbation budget at test time,
providing a defense thatdoes not rely on prior knowledge of the attack.
-
Quantified Performance Metrics: The system reports metrics like EERclean and EERadv across different detector settings and attacks, allowing researchers to quantify
clean condition discrimination and residual vulnerability after defense
while balancing the trade-off captured by the objective function.
Sources
- Securing Social Media Against Deepfakes using Identity, Behavioral, and Geometric Signatures
- Diffusion Reconstruction towards Generalizable Audio Deepfake Detection
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs