Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
cs.CR, cs.AI, cs.LG
Submitted: 2024-05-20
Updated: 2026-04-18
DOI: 10.1007/978-3-032-31583-0_5
Code: https://github.com/meta-llama/llama3
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Detecting Language Model Attacks with Perplexity
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
- COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
- Mistral 7B
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory
- Large Dual Encoders Are Generalizable Retrievers
- Training language models to follow instructions with human feedback
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Spectral Editing of Activations for Large Language Model Alignment
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Gemma: Open Models Based on Gemini Research and Technology
- LLaMA: Open and Efficient Foundation Language Models
- Finetuned Language Models Are Zero-Shot Learners
- Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs