Unread or Unenforced? Separating Representation from Enforcement Failure in Content Guards
cs.CR
Submitted: 2026-08-12
Updated: 2026-09-24
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
- The Llama 3 Herd of Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- LLM Safety From Within: Detecting Harmful Content with Internal Representations
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
- DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
- Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs