Eraser: Jailbreaking Defense in Large Language Models via Unlearning Harmful Knowledge
cs.CL
Submitted: 2024-04-08
Updated: 2026-09-20
Code: https://github.com/ZeroNLP/Eraser
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- PaLM 2 Technical Report
- Qwen Technical Report
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Toxicity in ChatGPT: Analyzing Persona-assigned Language Models
- Who's Harry Potter? Approximate Unlearning in LLMs
- LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
- LoRA: Low-Rank Adaptation of Large Language Models
- Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
- A Survey of Safety and Trustworthiness of Large Language Models through the Lens of Verification and Validation
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Certifying LLM Safety against Adversarial Prompting
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
- Making Harmful Behaviors Unlearnable for Large Language Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Baichuan 2: Open Large-scale Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering