AI Security Research Should Better Incentivize Defense Research
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "AI Security Research Should Better Incentivize Defense Research".
Nadia: This work examines an imbalance in artificial intelligence (AI) security research, finding that the field tends to produce more work on attacking AI systems than on defending them.
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: So, to start off, let's talk about the title itself, "AI Security Research Should Better Incentivize Defense Research," because it really frames the entire argument they are making about how we prioritize our efforts in this area.
Elias: The authors are Yiquan Zhang from The Hong Kong Polytechnic University and youqian.zhang@polyu.edu.hk, and they use this paper to argue that the current research landscape needs to shift its focus towards strengthening defensive work because of the existing imbalance between attacks and defenses in AI security research.
Priya: From a privacy standpoint, I'm interested in how this incentive structure might affect the kind of data we can actually secure; if defense gets more attention, maybe we see better protections for sensitive information being used by those large language models.
Nadia: That's a good point about the practical impact on privacy, Priya, and to expand on that imbalance, the paper points out that this isn't just a simple count issue but reflects how attack research is more readily published and rewarded than remediation efforts are.
Elias: They break down this imbalance by looking at subfields like federated learning and speech recognition, showing the asymmetry is not uniform across all topics.
The paper's summary: Nadia: Looking at the summary of "AI Security Research Should Better Incentivize Defense Research," it really highlights how this imbalance shows up in specific areas, showing that the issue isn't just general, but tied to certain attack classes.
Elias: They identify a distinct pattern where the most attack-heavy studies are all organized around a single attack class, such as black-box adversarial attacks or data inference privacy attacks against large language models.
Priya: That focus on specific offensive capabilities tells me that the literature is structured mostly around discovering and refining how AI systems can be broken rather than developing comprehensive defenses for a wide range of threats.
Nadia: Precisely, because they see this asymmetry in the SoK papers, they suggest that the most defense-heavy studies tend to emerge only when they focus on a very specific defense method rather than emerging in broad surveys of the overall threat landscape.
Elias: This suggests a structural dilemma where novelty is established more easily for attackers by simply identifying a new vector or applying an old technique to a new target, which sets a lower threshold for what counts as novel work in that domain.
The paper's improvements: Nadia: Now we get to the actual suggestions from the paper regarding how to fix this imbalance, and they propose several ways we should change our research incentives to encourage more defense work.
Elias: They suggest establishing "First-Class" Defense Contributions, meaning researchers should be rewarded primarily for developing robust, deployable defenses like prevention or certification rather than just demonstrating novel attacks.
Priya: From my perspective on measurement, I think standardizing evaluation criteria for defenses is crucial; we need universal benchmarks that test those defenses against adaptive adversaries across different models and real-world deployment conditions.
Nadia: That ties right into the evaluation asymmetry they discuss, because currently, an attack can succeed by breaking one version of a model under favorable conditions, whereas a defense is expected to work across many models and datasets.
Elias: They also suggest bridging that industry–academia gap by creating clear pathways from defensive ideas to actual deployment so that organizations can actually use the mitigation techniques being developed in the academic papers.
Conclusion: Nadia: So, to wrap up this discussion on "AI Security Research Should Better Incentivize Defense Research," the authors conclude that we need a better incentive structure focused on building and validating security solutions, not just identifying vulnerabilities.
Elias: They summarize that the bottleneck isn't just a lack of defense papers but also issues with how well defensive ideas are translated into systems organizations can actually use, stressing stronger public defense research and more realistic evaluations.
Priya: I think it’s important to remember that the paper flags a limitation: it relies on cited SoK papers and surveys, which means its analysis might not capture broader topics like safety or governance in the AI landscape.
Nadia: That's true, Priya; they admit they need a more comprehensive analysis by incorporating larger independent corpora in future work. But overall, the main point is that we need to improve how we reward and support work aimed at making AI systems fundamentally safer and more secure for deployment.
Youqian Zhang
The Hong Kong Polytechnic University
cs.CR, cs.AI
Submitted: 2026-05-22
Updated: 2026-09-30
Comments: 14 pages,3 figures,3 tables
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 80/100
The gist: This work examines an imbalance in artificial intelligence (AI) security research, finding that the field tends to produce more work on attacking AI systems than on defending them.
Key concepts
- Attack-Heavy SoKs
- These are knowledge papers where the focus is overwhelmingly on offensive methods. They tend to be organized around specific attack types, such as black-box attacks or data inference, rather than broad threat landscapes. This structure favors the discovery and characterization of new ways to break systems.
- Novelty Asymmetry
- Attack research can claim novelty easily by simply identifying a new attack vector or applying an old technique to a new target. In contrast, defenses must prove effectiveness across many models and datasets, making it structurally harder for defenses to establish novelty in the same way.
- Evaluation Asymmetry
- Attacks are judged by success against one version of a model under favorable conditions, whereas defenses are expected to work reliably across many models. Defenses are penalized for claiming universality because they must withstand attacks that succeed under specific, often non-universal, assumptions.
Terminology
Summary
This work examines an imbalance in artificial intelligence (AI) security research, finding that the field tends to produce more work on attacking AI systems than on defending them. This disparity suggests a structural lag in defensive progress, as attack papers are often evaluated under favorable conditions while defenses are held to a stricter standard. The paper argues that AI security research should better incentivize defense research to ensure systems are not only good at showing how they fail but also good at building and validating ways to make them secure.
The Core Finding on Imbalance
The analysis is based on a quantitative assessment of cited studies from 16 Systematization-of-Knowledge (SoK) papers and 21 survey papers, covering over 1,100 attack/defense-classified citations. The primary finding is that attack papers outnumber defense papers by 1.24:1
across the SoK corpus and by an average of 2.15
in the survey papers. The authors contend this imbalance is not merely a count but reflects a structural issue where attack research is more readily published and rewarded than remediation efforts.
Patterns of Asymmetry Across Subfields
The imbalance is not uniform across all topics; it varies substantially depending on the specific area of AI security being studied. The paper identifies two distinct patterns:
-
The most
attack-heavy SoKs are all organized around a specific attack class.
For instance, the five most skewed cases focus onblack-box adversarial attacks, data inference privacy attacks, membership inference against LLMs, neural network extraction via physical side channels, and anti-facial-recognition techniques.
In these areas, literature is structured primarily arounddiscovering, characterizing, and refining offensive capabilities.
-
The defense-heavy SoKs tend to emerge because they focus on a
specific defense method,
rather than being prominent in broad surveys of the threat landscape.
Structural Dilemmas for Defense Research
The paper outlines several dilemmas that contribute to the underincentivization of defensive work:
Novelty Asymmetry
:
Attack papers can establish novelty by identifying a new attack vector, applying a known technique to a new target, or showing that an existing system fails under a previously untested threat model.
This creates a structurally lower
threshold for novelty on the attack side.
Evaluation Asymmetry
:
Attacks and defenses are judged by different standards. An attack can succeed by breaking one version of one model under favorable conditions,
whereas a defense is expected to remain effective across many models and datasets. Defenses are penalized for incompleteness in a way attacks are not, because defensive claims are implicitly universal.
Publication Asymmetry
:
Attack papers often follow a structure that is easier to present and evaluate as a research contribution than incremental hardening (in defense research).
Defense claims rely on the more fragile assertion that a method withstands all evaluated attacks or mitigates a particular class of failures under specified assumptions.
The Industry–Academia Gap
A further factor contributing to the imbalance is the gap between academic findings and industrial practice. The authors note that researchers are often not aware of industry defense solutions
for certain model-extraction threat categories, suggesting that much of the most important work happens outside the kinds of outputs that paper counts can easily capture.
This results in a situation where attacks are often more visible because they reveal external vulnerabilities, while companies have limited incentive to publicize their mitigations in comparable detail.
Conclusion and Proposed Incentives
The authors conclude that AI security research should better incentivize defensive capacity in a broader sense.
They argue that the bottleneck is not solely a lack of defense papers but also issues with better translation of defensive ideas into systems that organizations can actually use.
Therefore, the necessary investment includes stronger public defense research, more realistic defense evaluation, clearer pathways from defensive ideas to deployment, and more institutional support for mitigation-oriented work.
The ultimate goal is for the field to be good at building, validating, and sharing ways to make them more secure and much safer.
Alternative Views Considered
The paper acknowledges alternative perspectives. One counterposition suggests that the imbalance might not be a problem if attack research is necessary to fully understand the main threats
before defenses can be reliably developed. Another view posits that the issue is not a lack of defense papers, but rather better translation,
meaning work focused on deployment and adoption is needed more than just publication counts. However, the authors maintain that without improving incentives around public defense work, this translational bottleneck alone is insufficient to fully resolve the imbalance.
Limitations
The analysis relies on cited SoK papers and surveys, which may not capture broader topics like safety or governance. Future work should aim for a more comprehensive analysis
by incorporating larger independent corpora and triangulating findings with a wider literature base. Furthermore, the paper identifies that specifying which interventions would mitigate this imbalance
remains an open question for future research.
Improvements for AI systems
Based on the provided paper, here are specific improvements that could be made to current AI security research and how these improvements would enhance AI systems:
- Improving Defense Research Incentives:
We must transition from merely observing an imbalance to actively correcting the incentive structure. This involves:
2.1. Establishing First-Class
Defense Contributions: Researchers should be explicitly rewarded for developing robust, deployable defenses (prevention, mitigation, certification) as primary scientific contributions, rather than solely for demonstrating novel attacks or vulnerabilities.
2.2. Standardizing Evaluation Criteria for Defenses: Develop rigorous, universal benchmarks that test defenses not just against isolated attacks but against adaptive adversaries across diverse models and real-world deployment conditions (e.g., testing robustness to batch size variations in FL or varying physical side-channel noise).
2.3. Bridging the Industry–Academia Gap: Create structured pathways for translating academic defensive findings into production-ready systems, including standardized engineering practices for hardening against identified vulnerabilities and sharing proprietary defense knowledge publicly where appropriate.
- Enhancing AI System Capabilities (The Result of Improved Incentives):
The improved research focus will lead to AI systems that are fundamentally more resilient and trustworthy:
-
Robustness Against Adversarial Inputs: Systems will be significantly harder to fool by malicious, imperceptible perturbations (e.g., in vision or language inputs) because defenses will be developed that are not brittle and can withstand
cherry-picked
threat models, leading to safer deployment in safety-critical applications like autonomous driving. -
Data Privacy and Model Integrity: AI models trained on sensitive data (like medical records or proprietary enterprise data) will have stronger inherent protections against inference attacks (Membership Inference Attacks) and model extraction, ensuring that sensitive training data remains private even when the model is deployed.
-
Trustworthy Deployment: Systems will move beyond mere accuracy metrics to include verifiable guarantees of performance under stress. This means AI systems can provide formal certification that they will operate reliably even when subjected to known classes of attacks or system manipulations, increasing their utility in high-stakes environments.
Abstract
This work examines an imbalance in artificial intelligence (AI) security research: the field tends to produce more work on attacking AI systems than on defending them. Drawing on related academic papers, we find biased attack-to-defense ratios across subfields, including federated learning, speech recognition, membership inference, large language models, etc. The imbalance possibly means far beyond a simple count: attack papers are routinely evaluated under favorable conditions that make threats look more severe than they are in practice, while defenses are held to a stricter standard that few can meet. The result is a literature rich in demonstrated vulnerabilities and thin on usable and deployed protections. We thus argue that AI security research should better incentivize defense research.
Sources
- SoK: Critical Evaluation of Quantum Machine Learning for Adversarial Robustness
- Intriguing properties of neural networks
- SoK: Evaluating Jailbreak Guardrails for Large Language Models
- Membership Inference Attacks on Large-Scale Models: A Survey
- Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs