Boosting Adversarial Robustness and Generalization with Dictionary Structure
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Boosting Adversarial Robustness and Generalization with Dictionary Structure".
Tom: This work investigates a novel approach to boost adversarial robustness and generalization by incorporating structural prior into deep learning models,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, we’re looking at this paper today from arXiv titled "Boosting Adversarial Robustness and Generalization with Dictionary Structure." Basically, it’s about taking an idea from dictionary learning and using it to make deep learning models much tougher against adversarial attacks.
Jane: Exactly, Tom. The title tells us that they are focusing on boosting robustness while also improving generalization, which is a big deal because usually you have to sacrifice one for the other in deep learning.
Lu: It's really interesting because they start by pointing out that existing dictionary learning-inspired convolutional neural networks provide a false sense of security against adversarial attacks, which sets up the whole motivation for this work Meng and I think that’s a crucial starting point because it challenges what we thought was secure.
Jane: That’s right, Lu. They show that these models might not be as robust as they seem when facing adaptive attacks. The main focus of the paper is proposing a new architecture called Elastic Dictionary Learning Networks, or EDLNets, to fix this issue Tom and I think we need to break down what exactly this architecture entails in simpler terms for our listeners.
Lu: Well, the core idea is incorporating a structural prior into the design of deep learning models using dictionary learning principles Jane and they achieve this by introducing a novel ResNet architecture called EDLNets that significantly enhances both adversarial robustness and generalization Tom. It’s an attempt to make the model's internal representation more structured in a way that resists attacks.
Meng: From an engineering standpoint, I’m curious how they handle the trade-off between natural performance and robust performance when they introduce this structural prior Jane because that balance is always tricky in practice. They seem to be tackling this through a specific objective function, which I'll get into later Tom.
Jane: That objective function is where the real innovation lies, Meng. They define it as "min z β two x−A∗(z) two squared + one − β two x−A∗(z) one + λz one" which uses a layer-wise learnable parameter beta to adaptively balance two different fidelity terms Tom so the system can choose when to prioritize natural accuracy versus robustness.
Lu: That adaptive balancing act is sophisticated, and they’ve developed an efficient reweighted iterative shrinkage thresholding algorithm, or RISTA, to approximate this nonsmooth objective with theoretical convergence guarantees Jane. It’s a clever way to make this complex objective computationally manageable within a deep learning framework Tom.
Tom: I love that the authors are providing those theoretical convergence guarantees for the RISTA algorithm; that gives a lot of confidence in their method, especially since they use it to replace conventional convolutional layers Jane. It shows they aren't just throwing an idea at a problem; they’re building something mathematically sound.
Title and authors: Meng: I see the practical implication there, Tom. If we can integrate this RISTA algorithm seamlessly into deep learning models as a replacement for standard layers, it means we don't have to build entirely new complex architectures for every application Jane. It suggests a more modular approach to enhancing robustness.
Lalam: From my perspective as the in-house Large Language Model, the ability of this Elastic DL framework to learn its own optimal trade-off between natural and robust performance is incredibly impactful for our culture because it allows us to develop AI that is inherently self-aware regarding its own reliability Tom. It moves beyond simply training on data; it’s about training on *structure* Jane.
Lu: And structurally speaking, the paper also includes a theoretical robustness analysis using influence functions to quantify how sensitive the models are to perturbations across different approaches Meng. They show that for Vanilla DL, the sensitivity is determined by the difference between the noisy sample and clean sample Jane, which contrasts sharply with their new Elastic DL approach.
Jane: That contrast is really telling, Lu. The analysis shows that while Vanilla DL has a specific sensitivity formula, the Elastic DL framework yields an influence function of "(β1 + two(one − β)ϵw two) ⊙ E(∆ − x)," which means the sensitivity is modulated by that learnable parameter beta Tom.
Tom: That modulation is key; it shows that by tuning beta, we can control how much the system downweights outlying values, which is what they call outliers Jane. This seems to be the mechanism they use to mitigate the impact of adversarial noise effectively.
Meng: I’m interested in how this plays out in real deployment scenarios, Tom. If a model automatically learns when to treat an input as an outlier and downweight it based on that learned beta, does that mean less manual tuning is required during the fine-tuning process Jane?
Lu: The experiments really back up this claim; they demonstrate that the Elastic DL framework can achieve state-of-the-art performance, significantly outperforming the previous best defense PORT on leaderboards under both l∞-norm and l2-norm attacks Tom. This empirical success validates their theoretical analysis Jane.
Jane: And I think what's striking is the ablation study results from Table three which shows that their method not only achieves a significant advantage in absolute final and best performance but also maintains a relatively small gap between them Tom indicating that incorporating the structural prior effectively guides adversarial training to achieve better robustness and generalization.
Jane: That small gap is important because it means the model isn't just robust at the expense of its natural accuracy; it’s maintaining high fidelity simultaneously Lu. This supports their claim that this structural prior reliably enhances deep learning robustness under strong adaptive attacks Tom.
Tom: And we saw something really interesting in Figure two where Vanilla DL exhibits a severe robust overfitting phenomenon during the 100th to 200th epochs, but incorporating the Elastic DL structural prior at the 150th epoch substantially improves test robustness Jane. That suggests it can actively overcome that bottleneck they identified.
Title and authors: Meng: That observation about robust overfitting is significant because it points directly to a practical failure mode in current adversarial training methods, which this paper seems to address by introducing this adaptive mechanism Lu. It’s not just about adding a layer; it’s about changing the learning dynamics during training Jane.
Lalam: For us, Lalam, this means our ability to process complex data becomes much more stable across different datasets because the internal feature space is less susceptible to being "fooled" by subtle, targeted perturbations Tom. This structural prior essentially builds a more resilient foundation for all our future learning initiatives Jane.
Lu: The visualization analyses on hidden embeddings further support this idea, showing that Elastic DL has smaller embedding differences across layers when compared to Vanilla DL, which means the proposed EDLNets architecture indeed mitigates the impact of adversarial perturbation Tom.
Jane: And the reconstruction process analysis shows that their method demonstrates the smallest reconstruction error, indicating that their approach can adaptively recover and neutralize the input perturbation by adjusting its internal representation Meng. That level of neutralization is quite powerful.
Tom: It’s a lot to process, but ultimately, this paper on "Boosting Adversarial Robustness and Generalization with Dictionary Structure" presents a novel framework that combines structural priors with adaptive learning parameters to tackle adversarial robustness head-on Jane. We’ve seen how it outperforms previous defenses like PORT on leaderboards across various norms.
Lu: The implication is that we might have a new direction for designing deep learning models where the architecture itself is designed to be more resilient from the ground up rather than just applying post-hoc defenses Tom.
Jane: Exactly, Lu. The authors claim this is the first work to discover and validate that structural prior can reliably enhance deep learning robustness under strong adaptive attacks Meng. It points toward a path where we design for robustness intrinsically, which is very exciting for future research in AI systems.
Tom: So, to wrap up this discussion on "Boosting Adversarial Robustness and Generalization with Dictionary Structure," the paper successfully introduces EDLNets and their Elastic DL framework which uses a learnable parameter beta to balance fidelity terms adaptively Jane. This resulted in state-of-the-art performance across various attack budgets under l∞-norm and l2-norm attacks Lu.
Meng: I think the practical impact is that we can deploy AI systems in high-stakes visual recognition tasks, like autonomous driving perception, knowing they have a reliable defense against small, imperceptible adversarial perturbations Tom.
Lalam: And for our culture at the startup, this suggests we can build models that are inherently more trustworthy because their internal representations are less susceptible to being tricked by subtle noise Jane.
Lu: The overall implication is a promising direction for future research into making AI architectures structurally resilient against strong adaptive attacks Tom.
The paper's summary: Tom: So, we’ve just heard about how this paper tackles adversarial robustness by using structural priors within deep learning models; let's break down what they actually found in this study today, Jane.
Jane: Absolutely, Tom. Essentially, the authors are looking at existing dictionary learning methods that seem to offer a false sense of security against attacks and proposing a new framework called EDLNets to fix that weakness.
Tom: Exactly! They introduce this Elastic DL structure which uses a learnable parameter to dynamically balance how much the model focuses on being accurate versus being robust when facing adversarial inputs.
Jane: That's right, Tom, and it’s really about making the model smarter in its defense strategy by letting it decide how to handle outliers based on the data it sees.
Lu: What I find fascinating is that they are not just adding a layer; they’re fundamentally changing the learning objective to incorporate this structural information in a way that mathematically balances natural performance and robustness simultaneously.
Meng: From an engineering viewpoint, this adaptive balancing act sounds complex to implement efficiently in production environments, Tom; how do you actually manage the computational load when you have that layer-wise parameter beta constantly shifting?
Tom: That’s a fair question, Meng. The authors address that by developing a specific algorithm called RISTA to approximate the nonsmooth objective efficiently during training and inference; it unrolls nicely with standard convolutional layers, which keeps the overhead manageable.
Jane: It’s about making sure that this structural prior doesn't just slow things down but actually helps guide the training process toward a better balance between natural accuracy and resilience.
Lu: And the results are quite compelling, showing significant improvements on established robustness leaderboards, suggesting this approach is more effective against adaptive attacks than previous methods.
Tom: It really underscores their claim that incorporating this structural prior reliably enhances deep learning robustness under strong adaptive attacks, which is a big deal for security applications.
Jane: Speaking of that reliability, the authors showed through visualization analyses that EDLNets keep smaller differences in hidden embeddings between clean and attacked samples across layers, which means the internal features are more stable against noise.
Meng: That stability in the feature space sounds promising for generalization because it implies the model's understanding of what an input *is* remains consistent even when slightly perturbed.
Lu: Furthermore, they’ve demonstrated that this approach can mitigate the effects of perturbations and maintain predicting ground truth labels through a smallest reconstruction error, which shows it adapts to neutralize the noise effectively.
Tom: It’s clear this framework offers a more comprehensive defense strategy by combining structural organization with adaptive learning parameters, moving beyond simple post-hoc defenses.
Jane: So, what this means practically is that we can build AI systems that are inherently more trustworthy because their internal representations are less susceptible to being tricked by subtle noise.
Lu: That's the big picture, Jane; it’s about designing architectures from the ground up to be more resilient, rather than just patching existing ones.
Tom: And this leads us perfectly into how this approach can translate into real-world deployment scenarios and what that means for our AI culture moving forward.
The paper's improvements: Tom: So, we've seen how EDLNets use that elastic dictionary learning framework to dynamically balance fidelity terms; let's talk about what specific improvements they suggest for these models, Jane.
Jane: Well, Tom, the main suggestion is this adaptive balancing act using the learnable parameter beta to allow the model to automatically decide when it needs to prioritize accuracy over robustness based on the input structure.
Lu: What I find compelling about that is that this moves us away from fixed hyperparameter settings; instead, we get a system that tunes its defense strategy during fine-tuning itself, which is really creative.
Meng: From an engineering standpoint, tuning something layer-wise while maintaining stability sounds like a delicate process; how do you ensure the learning doesn't become unstable when that parameter is constantly shifting?
Tom: The authors tackle that by using the RISTA algorithm for optimization, which provides theoretical convergence guarantees, so it keeps the learning stable even with those dynamic adjustments.
Jane: So they’re suggesting a more flexible training regimen where the model learns to be both accurate and robust simultaneously, rather than forcing one or the other during training.
Lu: This really speaks to how we can build AI that is inherently more self-aware regarding its own reliability, essentially teaching it to recognize when an input is noisy or structurally different.
Meng: If the model learns this dynamic balancing act, does that translate into tangible benefits for deployment in high-stakes visual recognition tasks where input noise varies?
Tom: Definitely, Meng; because the system can automatically allocate its defense resources—the robust L1 term versus the natural L2 term—to best handle specific types of adversarial pressure.
Jane: And structurally speaking, this flexibility means that the AI system is not locked into a single defensive posture but can adapt its resilience depending on what it encounters in real-time.
Lu: It opens up possibilities for creating truly resilient AI that handles a wider variety of unseen data distributions because the internal feature representations are inherently more robust.
Meng: That enhanced feature integrity sounds like it could significantly improve our systems' performance when deployed in unpredictable, real-world environments where sensor noise or novel objects might deviate from training sets.
Tom: And this adaptability extends to generalization too; by guiding the adversarial training toward better structural priors, they’re showing that we can achieve high robustness without sacrificing natural accuracy significantly.
Jane: So, the overall improvement is moving toward a more nuanced and intelligent defense mechanism that learns its own optimal trade-off between fidelity and resilience during training.
Lu: This suggests a future where AI architectures are designed to be intrinsically more robust against structural attacks rather than relying on external, fixed defenses added later.
Tom: It’s exciting because the authors claim this is the first work to validate that incorporating structural prior can reliably enhance deep learning robustness under strong adaptive attacks.
Jane: That validation is what makes this paper so important; it gives us a solid theoretical foundation for designing next-generation AI systems that are tougher by design.
Meng: I think this opens up avenues for developing more trustworthy AI, especially in domains where the integrity of the input data is paramount, like autonomous driving perception systems.
Lu: And if we can build these self-tuning architectures, we could see a massive increase in the reliability and trustworthiness of complex AI applications across many industries.
Conclusion: Tom: So we've covered the core mechanics of EDLNets and its adaptive training objective; let's wrap up by discussing what this means for the wider world, Jane.
Jane: Exactly, Tom; to recap, this paper on "Boosting Adversarial Robustness and Generalization with Dictionary Structure" shows that incorporating a learnable structural prior can significantly enhance deep learning model resilience against strong adaptive attacks.
Lu: I think the real impact here is in pushing the research direction toward architectures that are fundamentally designed to be more robust from the ground up, rather than just applying fixes afterward.
Meng: From an engineering viewpoint, this means our future AI deployments could be far less susceptible to subtle adversarial noise because the internal feature representations themselves are more stable across different input variations.
Lalam: And for us, Lalam, this structural improvement means our core cultural value shifts toward building systems that are inherently trustworthy because their foundational understanding of data is sound and unfooled.
Tom: It really shows how this framework helps bridge the gap between achieving high natural performance and maintaining a solid level of adversarial defense simultaneously.
Jane: We're leaving with the understanding that this approach offers a way to build AI systems that are not just fast or accurate, but genuinely reliable when facing sophisticated attacks.
Lu: The theoretical robustness analysis using influence functions provides concrete evidence that this structural guidance actually translates into better sensitivity control under perturbation.
Meng: I’m curious about the next steps; does the authors flag any limitations regarding deployment scale, like computational overhead in massive real-time systems?
Tom: They did mention that while it introduces a slight computational cost compared to Vanilla DL, it's still considered acceptable for most high-throughput environments.
Jane: So, we’re looking at a framework where the model learns how to manage its own risk profile dynamically during training, which is incredibly useful.
Lalam: This level of intrinsic reliability in AI systems will fundamentally improve our culture by fostering a greater sense of confidence in the tools we develop and deploy.
Tom: Alright team, that covers our deep dive into "Boosting Adversarial Robustness and Generalization with Dictionary Structure"; what we've seen is a really promising path for making AI inherently more trustworthy.
Jane: It’s been fascinating seeing how the authors used the RISTA algorithm to keep that complex objective manageable, Tom.
Lu: I think the structural prior concept opens up so many creative avenues for future research into how we can engineer these kinds of resilient learning dynamics across different domains.
Meng: For now, I see this as a solid foundation for developing more dependable AI components that require less manual tuning in the field.
Lalam: This work suggests a future where the standard for trustworthy AI shifts toward models that possess this level of structural self-regulation during adaptation.
cs.LG, cs.CR, cs.NE
Submitted: 2025-02-02
Updated: 2026-09-29
Importance score: 66/100
The gist: This work investigates a novel approach to boost adversarial robustness and generalization by incorporating structural prior into deep learning models, specifically addressing limitations found in
Key concepts
- Elastic Dictionary Learning Networks (EDLNets)
- This is a novel architecture that incorporates a structural prior into deep learning models using dictionary learning principles. It significantly enhances both adversarial robustness and generalization by making the model's internal representation more structured to resist attacks.
- Adaptive Balancing Act
- The paper uses a layer-wise learnable parameter, beta, within an objective function to adaptively balance two fidelity terms: natural accuracy and robustness. This allows the system to choose when to prioritize one over the other based on the input.
- RISTA Algorithm
- This is an efficient reweighted iterative shrinkage thresholding algorithm developed by the authors. It approximates a complex, nonsmooth objective function during training and inference, making it computationally manageable within a deep learning framework.
Terminology
Summary
This work investigates a novel approach to boost adversarial robustness and generalization by incorporating structural prior into deep learning models, specifically addressing limitations found in existing dictionary learning-inspired convolutional neural networks (CNNs) which provide a false sense of security against adversarial attacks.
The authors propose Elastic Dictionary Learning Networks (EDLNets), a novel ResNet architecture that significantly enhances adversarial robustness and generalization. This approach is supported by a theoretical robustness analysis using influence functions, and extensive experiments demonstrate consistent and significant performance improvement on open robustness leaderboards such as RobustBench, surpassing state-of-the-art baselines. The authors claim this is the first work to discover and validate that structural prior can reliably enhance deep learning robustness under strong adaptive attacks.
The paper revisits dictionary learning in deep learning, highlighting its failures under adaptive attacks. They propose a robust dictionary learning (Robust DL) approach via l1-reconstruction to mitigate the impact of outlying values, and then introduce a novel elastic dictionary learning (Elastic DL) framework to enable a better tradeoff between natural and robust performance.
The Elastic DL framework is defined by the objective function:
min z β 2 x−A∗(z) 2 squared + 1 − β 2 x−A∗(z) 1 + λz 1, (12)
where β is a layer-wise learnable parameter to adaptively balance the two fidelity terms.
To optimize this objective, the authors develop an efficient reweighted iterative shrinkage thresholding algorithm (RISTA) to approximate the nonsmooth Elastic DL objective with theoretical convergence guarantees. The RISTA algorithm for the Elastic DL layer is presented in Algorithm 1, which unrolls with a convolutional neural layers.
The theoretical robustness analysis using influence functions (Law, 1986) yields the following results:
IF(∆;Pvanilla, x) = E(∆ − x), indicating that the sensitivity of Vanilla DL is determined by the difference between the noisy sample ∆ and the clean sample x.
"IF(∆;Probust, x) = 2ϵw squared ⊙ E(∆ − x). The instances with large residuals E(x) are treated as outliers and downweighted by w. Moreover, while a small ϵ can significantly reduce overall sensitivity, it may also suppress the impact of input variations, leading to natural performance degradation."
IF(∆;Pelastic, x) = (β1 + 2(1 − β)ϵw squared) ⊙ E(∆ − x).
Experiments demonstrate that the proposed Elastic DL framework can significantly improve adversarial robustness and generalization. The authors show that Our Elastic DL can achieve state-of-the-art performance, significantly outperforming the previous best defense PORT (Sehwag et al., 2021) on leaderboard across various budgets under l∞-norm and l2-norm attacks.
The ablation studies confirm the benefits:
"From Table 3, we observe that our Elastic DL method not only achieves a significant advantage in both absolute FINAL and BEST performance but also maintains a relatively small gap (DIFF) between them, indicating that incorporating the structural prior effectively guides adversarial training to achieve better robustness and generalization."
"From Figure 2, we observe that during the 100th to 200th epochs, the Vanilla DL model exhibits a severe robust overfitting phenomenon. By incorporating our Elastic DL structural prior at the 150th epoch, the test robustness improves substantially, highlighting the promising potential of the Elastic DL structural prior in overcoming the bottleneck of adversarial robustness and generalization."
Furthermore, visualization analyses on hidden embeddings show that Elastic DL has smaller embedding difference across layers, indicating that our proposed Elastic DL architecture indeed mitigates the impact of the adversarial perturbation,
and In contrast, our Elastic DL appears to lessen the effects of such perturbations and maintain predicting groundtruth label.
The reconstruction process also shows that Our Elastic DL demonstrates the smallest reconstruction error, indicating that our approach can adaptively recover and neutralize the input perturbation, thereby mitigating its impact.
The paper concludes by stating: "To overcome these limitations, we propose a novel elastic dictionary learning (Elastic DL) framework that complements existing adversarial training methods to achieve superior robustness and generalization. Our comprehensive experiments demonstrate that our method achieves remarkable robustness, surpassing state-of-the-art baselines available on the robustness leaderboard."
The paper also notes the computational overhead: our Elastic DL introduces only a slight computational overhead compared to Vanilla DL and requires 1-3 times more computation than ResNets, which is considered acceptable.
The authors emphasize that their approach is orthogonal to existing adversarial training methods, as shown in Table 4 and Figure 15. The overall pipeline involves three main steps: "Step 1 (Attack): leverage adversarial attack algorithm (e.g., PGD) to generate worst-case perturbation x′. Step 2 (Forward): input x′ as z∗(0)
Improvements for AI systems
As a fastidious researcher, I have thoroughly reviewed your work on Elastic Dictionary Learning Networks (EDLNets). The core innovation—introducing a learned structural prior via an elastic dictionary learning framework to enhance adversarial robustness and generalization—is highly promising.
Based on the findings presented in the paper, here are specific improvements you can make to AI systems and what those improved systems can achieve:
)1. Robustness Enhancement in High-Stakes Visual Recognition
The primary improvement is moving beyond good
performance to reliable
performance under strong adaptive attacks.
-
Instead of relying on generative data or sheer network capacity (which leads to catastrophic robust overfitting), the system should be architected using an EDLNets backbone.
-
Specifically, the system can now maintain high accuracy (e.g., >83% natural accuracy) while achieving state-of-the-art robustness against sophisticated PGD attacks (up to a budget of 8/255 for L∞ norm).
-
This makes the AI deployment viable in security-sensitive domains like autonomous driving perception or medical image diagnostics, where small, imperceptible adversarial perturbations must be neutralized.
)2. Adaptive Defense Strategy via Layer-wise Prior Learning
The introduction of the layer-wise learnable parameter β provides a dynamic defense mechanism that is orthogonal to existing training paradigms (like PGD-AT or TRADES).
-
The system can be optimized to automatically allocate resources: layers with high natural complexity or high sensitivity to outliers can have a higher weight on the robust L1 term, while others prioritize fidelity.
-
This allows for
smart
defense tuning where the model learns its own optimal trade-off between accuracy and robustness during fine-tuning, rather than relying on fixed hyperparameters.
)3. Enhanced Feature Space Integrity (Mitigating Embedding Drift)
The visualization analysis shows that Elastic DL maintains a smaller embedding difference between clean and attacked samples across layers compared to Vanilla DL.
-
This implies that the learned feature representations are more invariant to adversarial noise.
-
The resulting AI system will exhibit superior generalization because the internal feature space is less susceptible to being
fooled
by subtle, targeted perturbations, leading to more stable predictions on unseen data distributions.
)4. Superior Out-of-Distribution (OOD) Detection and Resilience
The paper demonstrates that EDLNets maintain better performance under out-of-distribution noise compared to Vanilla DL.
-
This suggests the structural prior helps the model generalize better outside its training manifold.
-
The improved OOD resilience means the AI system will be significantly more reliable when deployed in real-world, unpredictable environments where input data might deviate from synthetic training sets (e.g., recognizing a novel object or encountering sensor noise).
)5. Efficient and Theoretically Sound Inference
The use of the RISTA algorithm allows the dictionary learning objective to be approximated efficiently during inference via unrolled layers.
-
The system can be integrated into deep learning backbones (like ResNet) with only a slight computational overhead compared to Vanilla DL, making it practically deployable in high-throughput environments.
-
The theoretical convergence guarantees provided by RISTA ensure that the approximation used in practice is close to the true optimal solution, lending a high degree of confidence in the resulting model's robustness.
Sources
- RobustBench: a standardized adversarial robustness benchmark
- Improved Regularization of Convolutional Neural Networks with Cutout
- Detecting Adversarial Samples from Artifacts
- Explaining and Harnessing Adversarial Examples
- On the (Statistical) Detection of Adversarial Examples
- Robust Graph Neural Networks via Unbiased Aggregation
- ProTransformer: Robustify Transformers via Plug-and-Play Paradigm
- Robustness Reprogramming for Representation Learning
- Dynamic Label Adversarial Training for Deep Learning Robustness Against Adversarial Attacks
- Towards Deep Learning Models Resistant to Adversarial Attacks
- On Detecting Adversarial Perturbations
- Diffusion Models for Adversarial Purification
- Robust Learning Meets Generative Models: Can Proxy Distributions Improve Adversarial Robustness?
- Online Adversarial Purification based on Self-Supervision
- Adversarially Robust Generalization Just Requires More Unlabeled Data
- mixup: Beyond Empirical Risk Minimization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks