GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
cs.CR
Submitted: 2025-09-29
Updated: 2026-08-28
Comments: Accepted by EMNLP 2026
Code: https://github.com/meta-llama/llama3
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
- OR-Bench: An Over-Refusal Benchmark for Large Language Models
- Privacy Side Channels in Machine Learning Systems
- MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
- AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
- Mistral 7B
- MCIP: Protecting MCP Safety via Model Contextual Integrity Protocol
- $R^2$-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning
- Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
- GuardReasoner: Towards Reasoning-based LLM Safeguards
- Ignore Previous Prompt: Attack Techniques For Language Models
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style Transfer
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs