CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization
Xinting Liao, Behnoosh Zamanlooy, Masoumeh Shafieinejad, David B. Emerson, Ruinan Jin, Deval Pandya, Xiaoxiao Li
cs.CR, cs.AI
Submitted: 2026-07-21
Comments: Accepted by COLM 2026 and AI4GOOD workshop
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- Quantifying LLM Biases Across Instruction Boundary in Mixed Question Forms
- Granite Guardian
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
- Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
- LiveBench: A Challenging, Contamination-Limited LLM Benchmark
- Black-box Optimization of LLM Outputs by Asking for Directions
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs