Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue
cs.CR, cs.AI, cs.IR, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning
- Gradient-Controlled Decoding: A Safety Guardrail for LLMs with Dual-Anchor Steering
- The Llama 3 Herd of Models
- Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
- Qwen2.5-Coder Technical Report
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs