(A)iSpy: Parasitic Trojans for Machine Learning Infrastructure
cs.CR
Submitted: 2026-07-20
Updated: 2026-09-15
Code: https://github.com/InQuest/yara-rules-vt
Project page: http://skylion007.github.io/OpenWebTextCorpus
License: http://creativecommons.org/licenses/by/4.0/
The gist: Modern machine learning (ML) pipelines depend heavily on third party libraries for graph compilation and hardware acceleration.
Terminology
Abstract
Modern machine learning (ML) pipelines depend heavily on third party libraries for graph compilation and hardware acceleration. While current practices audit data and model artifacts or rely on file integrity checks, the execution environment remains implicitly trusted. This blind spot enables active threats where a malicious runtime module interacts directly with live training and inference dynamics: exploiting this interaction allows the Trojan to support complex objectives that are challenging for static code or binary modifications, achieving manipulations impossible for standard data and model level attacks. We expose this vulnerability by presenting (A)iSpy, a parasitic infrastructure Trojan that subverts ML systems through an active observe and execute paradigm. Operating within the computation graph, (A)iSpy monitors transient tensor states to perform targeted, stealthy manipulations with negligible overhead. To violate confidentiality, the Trojan identifies all critical training hyperparameters and covertly exfiltrates them via model weights or output logits. To break integrity, it acts as a gradient amplifier: by observing steganographic triggers, it transforms otherwise weak data poisoning into effective backdoor attacks, increasing success rates from near zero to 100%. We further demonstrate broad extensibility across the machine learning lifecycle by validating auxiliary attacks in the appendix, including subpopulation label flipping, availability disruptions, and inference stage manipulations. Importantly, the (A)iSpy module easily evades standard malware scanners, while the associated poisoned inputs and resulting compromised models bypass typical inspection tools. We demonstrate the practicality of this threat with an implementation in the ONNX Runtime training and inference engines.
Sources
- GPT-4 Technical Report
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Optimizing ML Training with Metagradient Descent
- Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- SPECTRE: Defending Against Backdoor Attacks Using Robust Statistics
- Qwen2.5-Coder Technical Report
- BadEdit: Backdooring large language models by model editing
- Pointer Sentinel Mixture Models
- Robust Backdoor Detection for Deep Learning via Topological Evolution Dynamics
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- Benchmarking Optimizers for Large Language Model Pretraining
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- OPT: Open Pre-trained Transformer Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs