ZeroPur: Succinct Training-Free Adversarial Purification

arXiv:2406.03143 · cs.CV, cs.CR · Submitted 2026-08-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "ZeroPur: Succinct Training-Free Adversarial Purification".

Jane: The paper was written by Erhu Liu, Zonglin Yang, Bo Liu, Xianjia Meng, Xiuli Bi et al. from Chongqing University of Posts and Telecommunications and University of Nebraska-Lincoln and Northwest University and Northwestern Polytechnical University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone! Today we're diving into a paper that's been making waves in the AI security world, and it's called "ZeroPur: Succinct Training-Free Adversarial Purification."

Jane: And I have to say, Tom, the title alone got me excited. "Training-free" is a big deal. Usually, when you want to defend a model against adversarial attacks, you have to retrain it, or train some auxiliary network, which costs a ton of compute and time. This paper claims they can purify adversarial images without any of that.

Tom: Exactly, Jane. And that's the hook. The authors, Erhu Liu, Zonglin Yang, Bo Liu, and the team, they're basically saying, "Hey, we can clean up the mess an attacker makes without needing to teach the model anything new." That's a bold claim.

Jane: It is, but they back it up with a really elegant idea. They lean on the natural image manifold hypothesis. That's the idea that all natural, real-world images live on a kind of "surface" in the mathematical space the model sees. Adversarial attacks push images off that surface.

Tom: So, the attack is like knocking a marble off a table, and purification is trying to roll it back onto the table.

Jane: Precisely. And their method, ZeroPur, does this in two steps. First, they use a "Guided Shift" which uses a simple blurring operation to find a direction that points back toward that manifold. Then, they use "Adaptive Projection" to actually move the image back along that direction.

Tom: And the fact that they're using the victim classifier itself to do this, not an external generative model like a diffusion model, is what makes it so lightweight. It's a really clever repurposing of the model's own representations.

Jane: Right. And the implications are huge. Think about all the deployed systems out there running on edge devices. They can't afford to run a massive diffusion model to purify every single image. ZeroPur could be a practical, drop-in defense.

Tom: That's the dream, right? A defense that doesn't require you to change your infrastructure. I'm really curious to see how well it actually holds up against the strongest attacks, though. That's the real test.

Jane: Oh, we'll get to that. The paper has some impressive numbers on CIFAR and ImageNet. But let's first make sure everyone understands the core problem they're solving. It's a big one.

Tom: Definitely. So, adversarial attacks are these tiny, imperceptible tweaks to an image that cause a model to misclassify it. Like, a picture of a panda becomes a gibbon just by adding a little bit of noise. This paper is about cleaning that noise up.

Jane: And doing it without training. That's the "Zero" in ZeroPur. It's a succinct, elegant solution to a very messy problem.

Tom: I'm excited to dig into the "how" in the next segment. Stay tuned!

Summary: Tom: Welcome back. We're talking about "ZeroPur: Succinct Training-Free Adversarial Purification." Last segment we set the stage: it's a defense that doesn't need retraining. Now let's get into the nitty-gritty of how it actually works.

Jane: Right. So the paper's core insight is that a simple blurring operation provides a surprisingly good direction toward the natural image manifold. Think of it like this: if you have a noisy, adversarial image, blurring it a little bit gets you closer to a "clean" version of it, at least in the model's eyes.

Tom: And they call that first stage "Guided Shift." They iteratively move the adversarial image a tiny step toward its blurred counterpart, using the model's own feature embeddings as a guide. It's like taking small steps toward a goal, guided by a blurry map.

Jane: Exactly. But there's a catch. If you blur too much, you lose the content. If you blur too little, you don't get close enough to the manifold. So they found that after a few iterations, the process kind of stalls. The image gets closer, but it doesn't fully return.

Tom: That's where the second stage, "Adaptive Projection," comes in. It takes the direction found by Guided Shift and uses it as a momentum, but then it lets the image move more freely. It's like, "Okay, we know the general direction, now let's let the image find its own way back."

Jane: And they do this by maximizing the projection of the current image's features onto the direction established by the Guided Shift. It's a clever way to use the classifier's own understanding to pull the image back.

Tom: Now, the really clever part is how they handle natural images. If you apply this to a clean image, you don't want to move it off the manifold. So they add a perceptual regularization term. It's a constraint that keeps the purified image from deviating too much from the original input, which is crucial for maintaining standard accuracy.

Jane: And that's a big deal. A lot of defenses boost robustness but tank the model's performance on normal, clean images. ZeroPur tries to have both.

Tom: And the results are pretty wild. On CIFAR-ten with a ResNet-eighteen they get a robust accuracy of sixty-nine point six two percent against AutoAttack, which is a very strong attack. That's a huge jump from the twenty-nine point nine three percent baseline without defense.

Jane: It is. And on ImageNet, they're getting sixty-seven point zero nine percent robust accuracy, which is competitive with, and sometimes better than, methods that use heavy external diffusion models.

Tom: So the summary is: they've built a two-step, training-free pipeline that uses the victim model itself to pull adversarial images back to the natural manifold, and it works surprisingly well.

Jane: And it's not just about the numbers. It's about the philosophy. They're showing that you don't always need a bigger hammer. Sometimes, a simple, well-directed push is enough.

Tom: Let's talk about what this means for the field in the next segment. What are the real-world implications?

Improvements: Tom: We're back with "ZeroPur: Succinct Training-Free Adversarial Purification." We've covered the basics and the impressive results. Now, let's talk about what this actually improves and why it matters.

Jane: For me, the biggest improvement is the removal of the training burden. Look at the comparison tables in the paper. Methods like adversarial training require you to generate adversarial examples during training, which is computationally brutal. Other purification methods, like diffusion-based ones, require training or fine-tuning a massive generative model.

Tom: Right. And ZeroPur just... doesn't. It uses the off-the-shelf classifier. That's a paradigm shift. It means you can take any existing, already-deployed model and add this defense on top of it without touching its weights.

Jane: And that's not just a convenience. It's a practical enabler. Think about a medical imaging system that's been certified by regulators. You can't just retrain it because you found a new attack. ZeroPur could be a way to harden it without invalidating the certification.

Tom: That's a fantastic point. The paper also shows that ZeroPur is more robust to unseen attacks. Adversarial training often overfits to the specific attack used during training. ZeroPur doesn't have that problem because it's not learning a specific attack pattern.

Jane: Exactly. And they even test against adaptive attacks, where the attacker knows about the defense and tries to bypass it. They show that while an adaptive attack can reduce the benefit, it also makes the attack itself less effective. It's a trade-off that favors the defender.

Tom: And there's another improvement I want to highlight: the ablation study on the perceptual regularization. Without it, the standard accuracy drops significantly. With it, they maintain high performance on clean images while still getting the robustness boost. That balance is critical for real-world deployment.

Jane: Absolutely. And they also show the method is stable across different hyperparameters. You don't need to fine-tune the number of iterations for every single model. That's a sign of a robust, well-designed method.

Tom: So, the improvements are clear: it's cheaper, it's more general, it's more practical, and it doesn't sacrifice standard accuracy. This could be a real game-changer for how we think about defending AI systems.

Jane: It really could. And I think it opens up a new line of research. If a simple blur can guide us back to the manifold, what other simple, training-free operations could we use? The paper even mentions that replacing blur with total variation minimization improves performance further.

Tom: That's a great segue. Let's bring in our senior researcher, Lu, and our engineer, Meng, to get their take on this in the final segment.

Conclusion: Tom: We're wrapping up our discussion on "ZeroPur: Succinct Training-Free Adversarial Purification." It's been a fascinating deep dive.

Jane: It really has. And I want to bring in Lu and Meng to get their final thoughts. Lu, you're always thinking about the big picture. What's the most exciting implication for you?

Lu: For me, it's the democratization of defense. This method shows that you don't need a massive research lab with hundreds of GPUs to defend your models. The fact that it's training-free and uses the model itself means it can be applied anywhere, by anyone. That's a huge step toward making AI safer for everyone.

Meng: I agree with the accessibility, but as an engineer, I'm focused on the latency. The paper reports that ZeroPur runs at about two images per second on ImageNet with a ResNet-fifty on an A6000. That's a lot faster than diffusion-based methods, which can take minutes per image. But it's still not real-time for video.

Tom: That's a fair point, Meng. But for many applications, like verifying a single uploaded image or a batch of medical scans, a few seconds of processing is perfectly acceptable.

Meng: True. And the fact that it's a drop-in module with no training makes it easy to integrate into an existing pipeline. That's a big plus for us.

Jane: And I love that the paper is so transparent about its limitations. They acknowledge that it's less effective against very strong adaptive attacks like BPDA. That honesty is important for the field.

Lu: Absolutely. And that's what makes this a starting point, not an end point. It opens the door for future work on combining this lightweight approach with more powerful, but slower, methods. Maybe a hybrid system where ZeroPur handles the easy cases and a diffusion model is only called in for the hard ones.

Tom: That's a great vision. So, to summarize, "ZeroPur" is a training-free, two-stage adversarial purification method that uses a simple blur to guide adversarial images back to the natural image manifold. It's fast, it's practical, and it achieves state-of-the-art results.

Jane: And it challenges the assumption that you need heavy machinery to defend against adversarial attacks. Sometimes, a simple, clever push in the right direction is all you need.

Tom: Well said, Jane. We've covered the title, the summary, and the improvements. We've heard from our researcher and our engineer. It's time to say goodbye to this paper and get ready for the next one.

Jane: Thanks for joining us, everyone. We'll be back soon with another exciting paper from the arXiv. Until then, stay curious.

Tom: And stay safe out there. Goodbye!

Erhu Liu, Zonglin Yang, Bo Liu, Xianjia Meng, Xiuli Bi, Junwei Han, Bin Xiao

Chongqing University of Posts and Telecommunications · University of Nebraska-Lincoln · Northwest University · Northwestern Polytechnical University

cs.CV, cs.CR

Submitted: 2026-08-11

Comments: Accepted by IEEE Transactions on Image Processing (TIP) 2026

Code: https://github.com/erhul/ZeroPur

Project page: https://robustbench.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 42/100

The gist: Guided Shift (GS) and Adaptive Projection (AP).

Terminology

Summary

Summary

The paper introduces ZeroPur, a succinct, training-free adversarial purification method designed to defend deep neural networks (DNNs) against various unseen adversarial attacks without modifying the victim classifier. The method is grounded in the natural image manifold hypothesis, which posits that natural images reside on a specific manifold, and adversarial images are outliers of this manifold. Consequently, the purification process is framed as returning adversarial images to this manifold.

ZeroPur consists of two stages: Guided Shift (GS) and Adaptive Projection (AP). In the GS stage, adversarial examples are iteratively shifted toward their blurred counterparts to obtain embeddings closer to the natural image manifold. The paper observes that the blurring operator provides a reasonable direction towards natural image on the manifold for an adversarial sample. GS iteratively guides the adversarial sample toward its blurred counterparts, obtaining shifted embeddings that approximate the manifold. In the AP stage, a directional vector is constructed based on the shifted embedding, and the adversarial image is adaptively projected back onto the manifold. The paper states: AP constructs a directional vector by this shifted embedding to provide momentum, projecting adversarial images onto the manifold adaptively.

ZeroPur is independent of external models and requires no retraining of victim classifiers or auxiliary functions, relying solely on the victim classifier itself to achieve purification. The paper emphasizes: ZeroPur is independent of external models and requires no retraining of victim classifiers or auxiliary functions, relying solely on victim classifiers themselves to achieve purification.

The main contributions are:

  1. Analysis of adversarial attack and purification based on the natural image manifold hypothesis, showing that a simple blurring operator can bring adversarial examples closer to the manifold.

  2. Presentation of ZeroPur, a succinct adversarial purification approach with two stages (Guided Shift and Adaptive Projection), requiring no retraining of external models, parameterized functions, or the victim classifier.

  3. Extensive experiments demonstrating that the proposed approach outperforms most state-of-the-art auxiliary-based adversarial purification methods and achieves competitive performance compared to external model-based purification methods.

The paper provides a theoretical foundation via Theorem 1, which describes how perturbations are amplified layer by layer during forward propagation, causing adversarial embeddings to deviate from the natural image manifold. The theorem states: Suppose f: Rn → Rd is twice differentiable at point x... then we have the forward of f: f1...l (x + δ) = f1...l (x) + el (δ) where el (δ) can be viewed as the deviation distance from the natural image manifold.

The Guided Shift algorithm (Algorithm 1) initializes with noise, then iteratively applies blurring to generate manifold guidance, extracts embeddings, and purifies via gradient updates using cosine similarity as the distance metric. The Adaptive Projection algorithm (Algorithm 2) initializes with the input image, sets step size based on purification bound, and iteratively accumulates loss from candidate layers, incorporating perceptual regularization.

Experiments are conducted on three datasets (CIFAR-10, CIFAR-100, and ImageNet-1K) using various classifier architectures (ResNet-18, ResNet-50, WideResNet-28-10). The paper evaluates against four attacks: PGD, AutoAttack, DI2-FGSM, and BPDA. Results show that ZeroPur achieves state-of-the-art robust performance. For instance, on CIFAR-10 with ResNet-18, ZeroPur achieves 69.62% robust accuracy against AutoAttack (ϵ = 8/255), outperforming adversarial training and auxiliary-based purification methods. On ImageNet-1K, ZeroPur achieves 67.09% robust accuracy, significantly improving over existing methods.

The paper also addresses adaptive attacks where the adversary has full knowledge of the defense. The results show that while adaptive attacks can diminish the extra gain from ZeroPur, doing so sacrifices the original classification attack success rate, making it unworthy for the adversary to account for the defense.

Ablation studies validate the effectiveness of perceptual regularization, showing that incorporating it significantly improves standard accuracy without sacrificing robust accuracy. Sensitivity analysis examines the interaction between blur strength and training recipe, finding that classifiers trained with strong augmentation allow ZeroPur to benefit from stronger blur strength.

The paper concludes: ZeroPur underscores the promise of lightweight, training-free defenses for practical adversarial robustness. The source code is publicly available at https://github.com/erhul/ZeroPur.

Improvements for AI systems

Based on the scientific paper ZeroPur: Succinct Training-Free Adversarial Purification, here are the specific improvements I can make to an AI system and what the improved system can do:


1. Add a Training-Free Adversarial Purification Module (ZeroPur)

  • Implementation: Integrate a two-stage preprocessing pipeline before the classifier's forward pass:

  • Guided Shift (GS): Iteratively shift the input image toward its blurred counterpart in the feature space (using cosine similarity loss) for Tg=10 iterations with step size η1.

  • Adaptive Projection (AP): Project the shifted image back onto the natural image manifold by maximizing the dot product between the current feature difference and the guided feature difference across multiple candidate layers, with a perceptual regularization term (LPIPS-like) to prevent over-purification of natural images. Use Tp=50 iterations with step size η2=ϵpfy/Tp.

  • Key hyperparameters: Set purification bound ϵpfy = 1.25 × ϵatk (attack radius). For CIFAR-10/ResNet-18: λ1=1e-3, λ2=1e1; for CIFAR-10/WRN-28-10: λ1=1e-3, λ2=1e2; for CIFAR-100/ResNet-18: λ1=1e-4, λ2=1; for ImageNet/ResNet-50: λ1=1.5e-5, λ2=1.

  • Blur operator selection: Use Median filter (3×3) for classifiers trained without augmentation; use Gaussian blur (σ=1.2) for classifiers trained with standard or strong augmentation.

2. Enable Defense Against Unseen and Adaptive Attacks Without Retraining

  • Capability: The system can now defend against:

  • PGD, AutoAttack (l∞, ϵ=8/255 on CIFAR, 4/255 on ImageNet)

  • DI2-FGSM (blur-robust attacks)

  • BPDA (gradient approximation attacks) with robust accuracy remaining stable (e.g., 50.54% on WRN-28-10, 32.18% on ResNet-18 for PGD-40 on CIFAR-10)

  • No retraining required: The victim classifier, auxiliary functions, and external generative models remain untouched. This works with any off-the-shelf pretrained model (ResNet, WideResNet).

3. Improve Robust Accuracy Significantly

  • On CIFAR-10 (ResNet-18): Robust accuracy improves from 0% (undefended) to 69.62% under AutoAttack (ϵ=8/255), surpassing adversarial training methods like Gowal et al. (63.38%) and auxiliary-based methods like Mao et al. (67.15%).

  • On CIFAR-100 (ResNet-18): Robust accuracy reaches 41.42%, outperforming all compared AT/ABP methods (best prior: 34.37%).

  • On ImageNet-1K (ResNet-50): Robust accuracy reaches 67.09%, surpassing the best external model-based method (Mimicdiffusion: 62.16%) by 5%.

4. Maintain High Standard Accuracy

  • Trade-off control: The perceptual regularization term in AP ensures natural images are not over-purified. Standard accuracy remains high: 92.56% (CIFAR-10/ResNet-18), 91.81% (CIFAR-10/WRN-28-10), 71.88% (ImageNet-1K/ResNet-50).

5. Achieve Extreme Inference Efficiency

  • Throughput: The system processes images at 29.81 IPS (CIFAR-10/ResNet-18), 3.97 IPS (CIFAR-10/WRN-28-10), and 2.04 IPS (ImageNet-1K/ResNet-50) on a single RTX A6000 GPU. This is over two orders of magnitude faster than diffusion-based purification (e.g., Mimicdiffusion requires 156 seconds per ImageNet image on RTX 4090).

6. Add Robustness to Blur-Robust Attacks

  • Operator generalization: The system can swap the blur operator with Total Variation Minimization (TVM), improving robust accuracy against AutoAttack to 75.96% on CIFAR-10 (ResNet-18), demonstrating flexibility beyond blurring.

7. Provide Stable Performance Across Iteration Settings

  • Robustness to hyperparameters: The system maintains stable performance across a wide range of Tg (10–90) and Tp (10–90) settings, with robust accuracy varying by less than 1% (e.g., 69.25%–69.75% on CIFAR-10/ResNet-18).

  • Deploy in security-critical applications (e.g., autonomous driving, medical imaging, facial recognition) where adversarial attacks are a concern, without the computational cost of adversarial training or external generative models.

  • Defend against a wide range of unseen attacks (including adaptive and blur-robust attacks) with no additional training, making it practical for real-time or resource-constrained environments.

  • Maintain high accuracy on clean data while providing state-of-the-art robustness, avoiding the typical accuracy-robustness trade-off.

  • Work with any existing pretrained classifier (ResNet, WideResNet, etc.) without modifying its architecture or weights, enabling plug-and-play integration.

  • Operate at real-time speeds (e.g., 30 images per second on CIFAR-10) suitable for interactive systems, unlike slow diffusion-based defenses.

  • Provide interpretable purification: Visualizations show the system progressively removes adversarial noise while preserving semantic content, making it easier to audit and trust.

Sources

Related papers