Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks

arXiv:2609.05794 · cs.CR, cs.AI, cs.CL · Submitted 2026-09-05 · Read on arXiv

cs.CR, cs.AI, cs.CL

Submitted: 2026-09-05

Updated: 2026-09-05

Comments: 13 pages, 3 figures. Code: https://github.com/SparkShieldLab/bait-and-recover

Code: https://github.com/SparkShieldLab/bait-and-recover

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Related papers