Eraser: Jailbreaking Defense in Large Language Models via Unlearning Harmful Knowledge

arXiv:2404.05880 · cs.CL · Submitted 2024-04-08 · Read on arXiv

cs.CL

Submitted: 2024-04-08

Updated: 2026-09-20

Code: https://github.com/ZeroNLP/Eraser

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Sources

Related papers