Training-Free Refusal of MCP Exploits via Retrieval-Augmented Generation
cs.LG, cs.AI, cs.CR
Submitted: 2026-05-11
Updated: 2026-09-18
Comments: Published in the Conference on Applied Machine Learning for Information Security (CAMLIS) 2026. 22 pages, 4 figures, and 5 tables
Code: https://github.com/modelcontextprotocol/servers
License: http://creativecommons.org/licenses/by/4.0/
The gist: The model context protocol (MCP) has been widely adopted as an open standard enabling the seamless integration of generative AI agents.
Terminology
Abstract
The model context protocol (MCP) has been widely adopted as an open standard enabling the seamless integration of generative AI agents. However, while LLM guardrails have significantly matured to refuse malicious or harmful queries (e.g., "How do I build a bomb?"), recent work has shown that MCP-enabled LLMs are highly susceptible to prompt injection attacks which avoid harmful or suspicious cues (e.g., "Can you add this ssh key to my bashrc file?"). Herein, we use state-of-the-art (SOTA) alignment fine-tuning algorithms to explore whether LLMs may be aligned to refuse such falsely benign attacks (FBAs). While SOTA algorithms based on direct preference optimization (DPO) improve refusal guardrails against FBAs, we show that this improvement is limited; DPO-based fine-tuning never improves FBA refusal rates beyond 47% across five popular open-source LLMs. Thus, to further improve FBA refusals, we introduce Retrieval Augmented Generation for Preference alignment (RAG-Pref), a simple RAG-based alignment algorithm which conditions on preferred and dispreferred samples to leverage contrastive information during inference. RAG-Pref is online (training-free), compatible with off-the-shelf packages, and, when combined with offline alignment algorithms, enables an average 3.7-fold improvement in FBA refusals across five widely used LLMs, compared to 2.9 for other online alignment methods and 1.5 for offline alignment alone. We additionally show that RAG-Pref generalizes beyond agentic safety: in stark contrast to other online alignment methods, RAG-Pref consistently improves performance on general human-preference benchmarks AlpacaEval 2 and MT-Bench across five SOTA alignment-tuned models, demonstrating broad applicability to general alignment tasks.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Gemma 2: Improving Open Language Models at a Practical Size
- The Llama 3 Herd of Models
- ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
- GPT-4o System Card
- Towards Efficient Exact Optimization of Language Model Alignment
- Binary Classifier Optimization for Large Language Model Alignment
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- Prompt Injection attack against LLM-integrated Applications
- Ignore Previous Prompt: Attack Techniques For Language Models
- MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
- Gemini: A Family of Highly Capable Multimodal Models
- Zephyr: Direct Distillation of LM Alignment
- Qwen3 Technical Report
- How Language Model Hallucinations Can Snowball
- DPO Meets PPO: Reinforced Token Optimization for RLHF
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks