Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning
cs.LG, cs.CR
Submitted: 2026-09-27
Updated: 2026-09-29
Terminology
Sources
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Learning to Inject: Automated Prompt Injection via Reinforcement Learning
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
- EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- Muse Spark Safety & Preparedness Report
- Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
- Ignore Previous Prompt: Attack Techniques For Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming
- RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
- Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
- Qwen3 Technical Report
- CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks