Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
Maria Rosaria Briglia, Igor Maljkovic, Antonio Emanuele Cinà, Luca Oneto, Iacopo Masi, Fabio Roli
cs.CR, cs.AI, cs.LG
Submitted: 2026-07-28
Code: https://github.com/trusted-user/qwen3-vlfast-inference
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Architectural Backdoors in Deep Learning: A Survey of Vulnerabilities, Detection, and Defense
- Qwen3-VL Technical Report
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
- Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- HOD: A Benchmark Dataset for Harmful Object Detection
- DiffGuard: Text-Based Safety Checker for Diffusion Models
- The Llama 3 Herd of Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs