NeuPerm: Disrupting Malware Hidden in Neural Network Parameters by Leveraging Permutation Symmetry

arXiv:2510.20367 · cs.CR · Submitted 2025-10-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "NeuPerm: Disrupting Malware Hidden in Neural Network Parameters by Leveraging Permutation Symmetry".

Elias: Pretrained deep learning model sharing exposes end-users to cyber threats where attackers hide self-executing malware inside neural network parameters, and this work proposes NeuPerm,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So we're discussing the paper "NeuPerm: Disrupting Malware Hidden in Neural Network Parameters by Leveraging Permutation Symmetry." The core idea seems to be a zero-trust technique against malicious steganography hidden inside deep learning models that are shared between researchers and users. What does this paper actually claim is the main problem they are addressing?

Elias: Well, the thesis of NeuPerm is that attackers can embed self-executing malware into neural network parameters, creating these things called stegomalwares, which pose a serious threat to the ML supply chain because they can spread easily and don't degrade model performance forty-one, fifteen, fourteen <ref:2510.20367#pg1,threat to the ML supply chain>.

Priya: From a privacy standpoint, what concerns me is how these models are being shared; it sounds like the potential for widespread distribution of something malicious is high. I wonder if there are any specific types of data or model architectures that are particularly vulnerable to this kind of parameter hiding.

Nadia: Exactly, Priya, and the paper points out that neural networks have a much greater capacity for hidden data than traditional media because the models themselves can weigh multiple gigabytes; for example, the Llama3 point 3 70B LLM weighs approximately one hundred forty gigabytes forty-one <ref:2510.20367#pg1,neural networks have a much greater capacity for hidden data>.

Elias: That size difference is huge when you consider how easily these stegomalwares can propagate, potentially avoiding detection by common anti-virus software and malware detection systems fourteen, forty-five <ref:2510.20367#pg1,detection by common anti-virus software>.

Priya: So, the concern isn't just the presence of data, but the fact that this hidden malicious code operates within a structure that is designed to be trusted for its computational function. What is their proposed solution to stop this kind of embedding?

Nadia: The paper introduces NeuPerm as a simple yet effective way to disrupt these attacks by exploiting the theoretical property of neural network permutation symmetry, which they claim has little to no effect on model performance <ref:2510.20367#pg0>.

Elias: That reliance on permutation symmetry is interesting; it suggests that swapping units in a hidden layer doesn't change the computation itself, but it changes the parameter matrices, and that's where NeuPerm steps in to disrupt the attack <ref:2510.20367#pg2>.

Paper summary: Priya: If this symmetry holds true across different architectures, does this mean we could build a more generalized defense mechanism against these kinds of hidden payloads? I'm thinking about how robust the method is against different network types.

Nadia: The researchers are specifically looking at how to apply these permutations based on the architecture, for instance, they perform operations like "W1,b1 = W1PI, b1PI" in FC-FC blocks <ref:2510.20367#pg0>.

Elias: They also adapt their approach for different types of layers; for example, in CONV-CONV or CONV-BN-CONV blocks, they permute the first axis of the first CONV’s parameter matrices and the second axis of the second CONV’s parameter matrix <ref:2510.20367#pg0>.

Priya: And for attention mechanisms, like ATTN blocks, they adapt by permuting "the heads," which they adjust based on Grouped Query Attention by shuffling query, key, value, and projection parameter matrices <ref:2510.20367#pg0>.

Nadia: So the paper lays out a concrete way to apply these theoretical symmetries across various common neural network types to actively disrupt the hidden payload without altering the model's intended function <ref:2510.20367#pg1>.

Elias: The analysis shows that for attacks lacking error correction, such as StegoNet or EvilModel, the success probability is bounded by d L, meaning for any d less than one, this probability is almost surely zero <ref:2510.20367#pg1>.

Priya: That's a strong statement about disruption; it suggests that against simpler embedding techniques, the chance of extracting the payload is essentially zero if the perturbation parameter d stays below one. But what happens when we look at more sophisticated attacks?

Nadia: The paper addresses that with error-resilient attacks like MaleficNet fifteen, which use error-correcting codes, where the success probability is bounded by a more complex expression involving an additive Hoeffding bound <ref:2510.20367#pg1>.

Elias: And the empirical results they present are quite telling; for CNN and LLM MaleficNet stegomodels, full NeuPerm caused the Signal-to-Noise Ratio to become negative, which strongly asserts that the attack was disrupted and the payload cannot be extracted <ref:2510.20367#pg1>.

Priya: That is a very concrete result; it moves beyond just theoretical bounds and shows practical disruption in real models like DenseNet121, ResNet50/one hundred one and VGG11 on the CNN side <ref:2510.20367#pg2>.

Nadia: And they've applied it to LLMs too, specifically mentioning Llama-three point two-1B alongside CNNs <ref:2510.20367#pg2>, which shows the scope of this technique is quite broad across different model sizes and types.

Paper summary: Elias: Comparing NeuPerm against other methods like adding random noise or quantization reveals that NeuPerm is quick because it only requires reordering the axes of the parameter matrices in place, and it largely does not degrade performance thanks to that symmetry property <ref:2510.20367#pg0>.

Priya: That’s a key distinction we need to track: while noise can degrade performance, NeuPerm avoids needing retraining or fine-tuning for recovery, which speaks to its practical utility in a deployment setting.

Nadia: So the paper presents NeuPerm as quick and generic because it doesn't require that kind of extensive effort from the end user, contrasting with methods like quantization that need retraining <ref:2510.20367#pg0>.

Elias: The implication here is that we might be looking at a way to secure model sharing without imposing heavy computational overhead on the people using those models, provided the permutation symmetry holds as assumed <ref:2510.20367#pg2>.

Priya: What about the limitations they acknowledge? They state clearly that NeuPerm is only applicable to neural networks that actually possess this permutation symmetry property, suggesting future work might need to extend this methodology to other types of symmetry <ref:2510.20367#pg2>.

Nadia: That limitation tells us exactly where the research needs to go next, focusing on identifying which network structures truly exhibit these symmetries for this type of defense <ref:2510.20367#pg1>.

Elias: The authors also noted that if the payload is too large for a given architecture, they indicate that a "-" sign means the payload is too big, which sets a practical boundary on what this specific application can handle <ref:2510.20367#pg2>.

Priya: So, to summarize our discussion on "NeuPerm: Disrupting Malware Hidden in Neural Network Parameters by Leveraging Permutation Symmetry," the paper introduces a method that uses permutation symmetry to disrupt steganography without performance degradation, showing success against attacks like MaleficNet.

Nadia: And this points toward a future where securing model sharing might involve simple, in-place parameter reordering rather than complex retraining processes.

Elias: It suggests that the theoretical property of permutation symmetry provides a viable avenue for zero-trust defense in this context, contingent on the accuracy of those underlying assumptions <ref:2510.20367#pg2>.

Priya: The impact could be significant for anyone working with pre-trained models, as it offers a straightforward mechanism to counter a sophisticated threat vector that has been quite hard to detect previously.

Conclusion: Nadia: So, we've seen how NeuPerm uses permutation symmetry to disrupt hidden malware in neural network parameters, and now we need to talk about what this paper is actually called and who wrote it. Elias, can you tell us a bit about the title and the authors of "NeuPerm: Disrupting Malware Hidden in Neural Network Parameters by Leveraging Permutation Symmetry"?

Elias: I can confirm that the title clearly states the mechanism being used, focusing on permutation symmetry to disrupt malware embedded in neural network parameters. The authors are Samuel K. Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa; they're researchers who have done a lot of work in this area before.

Priya: From my perspective as someone focused on privacy and measurement, I'm curious about what these authors were aiming to achieve by focusing specifically on permutation symmetry as the core defense mechanism in that title. What’s the fundamental concept behind that specific choice?

Nadia: Well, Priya, it means they aren't just slapping some random noise on the data; they're targeting a structural property of the neural network itself—that underlying symmetry—to create a zero-trust barrier against these hidden attacks. It’s about using the network's own mathematical structure against the steganography.

Elias: Exactly, Nadia, and from a cryptographic standpoint, that implies they are exploiting an inherent redundancy in how certain layers process information; if you can permute those units and the function stays the same, that’s where we find our leverage. It suggests a very targeted attack surface for disruption.

Priya: I see why that matters for measurement research because it points toward a defense that works on the model's architecture rather than just treating the input data as a black box; it’s leveraging known properties of deep learning structures. Does this symmetry hold up across different types of networks, or is it very specific?

Nadia: That’s a key point for me, Priya; if it only works on certain architectures, then its practical application is limited to those specific models. The authors do flag that NeuPerm is only applicable to neural networks that possess this permutation symmetry property.

Elias: They are honest about the limitations there; they've set a clear boundary by stating it doesn't work everywhere, which is important for anyone trying to implement this in a real-world scenario. It tells us we need more research into identifying those specific network structures that have this property.

Priya: So, the implication is that while the concept is powerful—using symmetry for defense—the next step isn't just applying it broadly; it’s about understanding precisely which network designs are vulnerable to this kind of parameter hiding and which ones offer the necessary symmetry for NeuPerm to function.

Nadia: Precisely, Priya, so we move from proving a method works against existing attacks to figuring out where we can deploy this defense most effectively across different model types. It’s about moving from a theoretical proof to practical deployment strategy. **Show End**

Ariel Cyber Innovation Center · Ariel University

cs.CR

Submitted: 2025-10-23

Updated: 2026-10-07

Code: https://github.com/danigil/NeuPerm

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 73/100

The gist: Pretrained deep learning model sharing exposes end-users to cyber threats where attackers hide self-executing malware inside neural network parameters, and this work proposes NeuPerm, a simple yet

Key concepts

Neural Network Permutation Symmetry
This is a theoretical property where rearranging the order of neurons within a hidden layer in a neural network does not change the function that the network computes. NeuPerm exploits this property to disrupt hidden data, as attackers often rely on fixed positions for their malware.
Steganography Attacks (e.g., MaleficNet)
These are methods where malicious code is hidden inside the weights or parameters of a neural network, making the model appear normal while carrying a payload. Past defenses often struggle against these because they are designed to be error-resilient and hard to detect.
Zero-Trust Technique
NeuPerm acts as a zero-trust defense by assuming that any hidden data within the model parameters is potentially malicious. It actively intervenes by applying structural permutations to neutralize the hidden payload, ensuring trust is not implicitly granted to the model's internal state.

Terminology

Summary

Pretrained deep learning model sharing exposes end-users to cyber threats where attackers hide self-executing malware inside neural network parameters, and this work proposes NeuPerm, a simple yet effective zero-trust technique leveraging permutation symmetry to disrupt such attacks with little to no effect on model performance.

The gist

NeuPerm is a simple yet effective zero-trust technique for disrupting neural network steganography by leveraging the permutation symmetry of neural networks, showing success against state-of-the-art attacks like MaleficNet.

Threat Model and Motivation

The research addresses the ML supply chain attack vector where adversaries upload models infected with malware to Pretrained Model (PTM) hubs. Past research has shown attackers using datahiding (steganography) techniques to hide malware inside neural network parameters, resulting in stegomalwares that pose a security threat because they can be easily propagated and do not degrade model performance. While past countermeasures like quantization or fine-tuning exist, they are complex, require expertise, and are often ineffective against resilient attacks like MaleficNet. NeuPerm is introduced as a method that is simple, yet effective zero-trust technique for disrupting neural network steganography that leverages the permutation symmetry of neural networks.

Mechanism of NeuPerm

NeuPerm operates by exploiting the theoretical property of neural network permutation symmetry, which holds whenever neurons inside a hidden layer can be permuted and the neural network still computes the same function. The implementation involves applying permutations to specific parameter matrices based on the architecture:

  1. For FC-FC blocks, it performs operations like W1,b1 = W1[PI], b1[PI] and W2 = W2[:,PI].

  2. For CONV-CONV/CONV-BN-CONV blocks, it permutes the first axis of the first CONV’s parameter matrices and the second axis of the second CONV’s parameter matrix.

  3. For ATTN blocks, it permutes the heads, adapting to Grouped Query Attention (GQA) by permuting query, key, value, and projection parameter matrices.

Effectiveness Against Steganography

The effectiveness of NeuPerm is analyzed through a security game where the adversary attempts to extract the payload after the defender applies a permutation. The analysis shows that for attacks with no error correction capabilities (like StegoNet or EvilModel), the success probability S is bounded by d L, where d is related to the maximum output dimension, meaning "for any d < 1, this probability is a.a.s. 0." For error-resilient attacks like MaleficNet [15], which use error-correcting codes, the success probability S can be bounded by a more complex expression involving an additive Hoeffding bound:

S ≤ d L + exp−2 [(1 − δ) − d]2 L, where δ is the error correction capability. The empirical results demonstrate that full NeuPerm caused the Signal-to-Noise Ratio (SNR) to become negative for CNN and LLM MaleficNet stegomodels, which strongly asserts that the attack was disrupted and the payload can not be extracted.

Comparative Analysis of Disruption Methods

NeuPerm is compared against other disruption techniques, including adding random normal noise, fine-tuning, parameter pruning, and quantization. The comparison reveals key distinctions:

- Quick Generic Does not Require Retraining Does not Degrade Performance Effectiveness

Adding random noise is the simplest method but can degrade performance as empirical results show. Fine-tuning is not quick and not generic. Parameter pruning and quantization both require fine-tuning to recover lost performance. NeuPerm, conversely, is described as being quick, as it only requires reordering the axes of the model parameter matrices in place, and it does not require fine-tuning, and largely does not degrade model performance thanks to the permutation symmetry property.

Conclusion on Practicality

The paper concludes that NeuPerm is a significant advancement because it disrupts resilient error-correcting steganography like MaleficNet, which is only achieved otherwise (presumably) using quantization, a complex method that requires retraining, expertise, and resources. It is the first work to develop a method that explicitly disrupts MaleficNet, and it has been successfully applied to CNNs (DenseNet121, ResNet50/101, VGG11) and LLMs (Llama-3.2-1B). The paper notes that NeuPerm is only applicable to neural networks that possess the permutation symmetry property, suggesting future work may extend this methodology to other types of symmetry. It is also noted that "a "- sign means the payload is too large for the given architecture.

References

[1] Samuel K. Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the NeuPerm research, along with what those improved systems can do:


  1. The ability to detect and neutralize hidden, self-executing malware embedded within pre-trained deep learning models (stegomalware).

  2. The creation of a robust defense mechanism against sophisticated steganography techniques like Least Significant Bit (LSB) substitution, sign-mapping, and value-mapping attacks.

  3. Enhancing the security of the ML supply chain by verifying that shared or pre-trained models have not been tampered with by malicious actors before deployment to end-users.

  4. Developing a zero-trust disruption technique for neural networks that does not require complex, resource-intensive processes like full fine-tuning, retraining, or quantization (which often degrade performance).

The improved AI system can perform the following specific functions:

  1. A pre-trained Large Language Model (LLM) deployed by an enterprise or researcher can be scanned using NeuPerm to check for hidden malicious payloads embedded in its weights.

  2. If NeuPerm detects a payload, the system automatically disrupts the model's internal structure by applying a random permutation of hidden layer units, effectively scrambling the malware and rendering it inert.

  3. The system can verify that models downloaded from public hubs (like HuggingFace) are safe against steganographic attacks before they are fine-tuned or used in sensitive applications (e.g., medical diagnosis or finance).

  4. The resulting AI application maintains its original high performance (with negligible accuracy drop, up to 0.01% for CNNs and near baseline for LLMs), ensuring that the security measure does not compromise the AI's utility or accuracy during inference.

Abstract

Pretrained deep learning model sharing holds tremendous value for researchers and enterprises alike. It allows them to apply deep learning by fine-tuning models at a fraction of the cost of training a brand-new model. However, model sharing exposes end-users to cyber threats that leverage the models for malicious purposes. Attackers can use model sharing by hiding self-executing malware inside neural network parameters and then distributing them for unsuspecting users to unknowingly directly execute them, or indirectly as a dependency in another software. In this work, we propose NeuPerm, a simple yet effec- tive way of disrupting such malware by leveraging the theoretical property of neural network permutation symmetry. Our method has little to no effect on model performance at all, and we empirically show it successfully disrupts state-of-the-art attacks that were only previously addressed using quantization, a highly complex process. NeuPerm is shown to work on LLMs, a feat that no other previous similar works have achieved. The source code is available at https://github.com/danigil/NeuPerm.git.

Sources

Related papers