Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents
Mingxiao Liu, Yitong Li, Haoren Zhao, Yaoxiang Bian, Jianan Ma, Jian Zhang, Jialuo Chen, Xinhao Deng, Zhen Wang
Hangzhou Dianzi University · Hangzhou Dianzi University · Hangzhou Dianzi University · Hangzhou Dianzi University · Ant Group · Hangzhou Dianzi University · Ant Group · Ant Group · Hangzhou Dianzi University
cs.CR
Submitted: 2026-07-30
Comments: 19 pages, 8 figures, The code is publicly available at https://github.com/Limax666/AudioAgentSecurity
Code: https://github.com/Limax666/AudioAgentSecurity
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Empirical Analysis of Large Vision-Language Models against Goal Hijacking via Visual Prompt Injection
- WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
- Gemini: A Family of Highly Capable Multimodal Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment through Latent Acoustic Pattern Triggers
- Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
- Qwen3-Omni Technical Report
- Qwen3 Technical Report
- AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
- AutoGLM: Autonomous Foundation Agents for GUIs
- Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection
- MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
- GPT-4o System Card
- Step-Audio 2 Technical Report
- Prompt Injection Attack to Tool Selection in LLM Agents
- PromptLocate: Localizing Prompt Injection Attacks
- Beyond Pipelines: A Survey of the Paradigm Shift toward Model-Native Agentic AI
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- Toward Efficient Agents: Memory, Tool learning, and Planning
- NoiseAttack: An Evasive Sample-Specific Multi-Targeted Backdoor Attack Through White Gaussian Noise
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs