Geometric Configurations of Perturbed Jailbreak Prompts

arXiv:2607.20581 · cs.CR, cs.AI · Submitted 2026-07-22 · Read on arXiv

Lynn Delcon, Andres Algaba, Vincent Ginis

cs.CR, cs.AI

Submitted: 2026-07-22

Comments: 21 pages, 9 figures, 7 tables, 2nd Workshop on Safe AI (SafeAI)

Code: https://github.com/elder-plinius/L1B3RT4S

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety.

Terminology

Abstract

Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper, we investigate the internal representations of such string-level perturbed jailbreak inputs in the small weight models of the Qwen-2.5-1.5B/-3B/-7B-Instruct and Llama-3.2-1B/-3B/-3.1-8B-Instruct families. We select two representation spaces: the last-layer-last-token embedding space and the top-50 next-token probability space. The former space separates prompts based on their spelling and format, while the latter space is effectively one-dimensional but appears more complex to cluster. Within our refusal-dominated answer set we find no behavioral hyperplane in either space. Only the next token "Sure" in the 1.5B Qwen model, and both tokens "," and ".C.C" in the 1 Llama model, display a significant association with a compliant-labeled answer.

Sources

Related papers