From ASR to ASP: Evaluating Prompt Attack Vulnerabilities Against Open-Source LLMs
cs.CR, cs.CL
Submitted: 2025-05-20
Updated: 2026-09-25
Code: https://github.com/werywjw/Prompt-Injection
Terminology
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Stable LM 2 1.6B Technical Report
- Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
- GPT-4o System Card
- Prompt Injection attack against LLM-integrated Applications
- Mistral 7B
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- The Llama 3 Herd of Models
- Gemma: Open Models Based on Gemini Research and Technology
- Ignore Previous Prompt: Attack Techniques For Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- SOS! Soft Prompt Attack Against Open-Source Large Language Models
- Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs