FATS: A Prompt Injection Attack Utilizing Feign Security Agents with Deceptive Few-shots Learning
cs.CR, cs.AI
Submitted: 2024-10-11
Updated: 2026-09-05
Comments: The final version
DOI: 10.1007/978-981-95-7078-2_47
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Breaking Down the Defenses: A Comparative Survey of Attacks on Large Language Models
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs