RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
Benyamin Tafreshian, Prathamesh Dhake
cs.CR, cs.AI
Submitted: 2026-07-29
Comments: This manuscript supersedes the preliminary version available as arXiv:2511.18790. The work has been substantially revised, expanded, and reorganized, with a refined threat model, revised methodology, clearer stage-level evaluation criteria, and expanded analysis of moderation bypass, instruction reconstruction, and execution
Code: https://github.com/btafreshian/AIsecStudy
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) are becoming increasingly integrated into mainstream development platforms and daily technological workflows, typically behind moderation and safety controls.
Terminology
Abstract
Large language models (LLMs) are becoming increasingly integrated into mainstream development platforms and daily technological workflows, typically behind moderation and safety controls. Despite these controls, preventing prompt-based policy evasion remains challenging, and adversaries continue to "jailbreak" LLMs by crafting prompts that circumvent implemented safety mechanisms. Prior work has established cipher-mediated interaction, code-embedded decryption, prompt decomposition and reconstruction, and layered custom encryption as viable attack primitives. However, reported evaluations generally collapse visible acceptance, successful recovery of the concealed request, and subsequent execution into an aggregate attack-success outcome. This leaves limited evidence about where multistage prompt-transformation attacks fail within an observable black-box interaction. This paper introduces RoguePrompt, a jailbreak pipeline that partitions a forbidden prompt and applies two nested encodings, Vigenere followed by ROT13, along with natural-language reconstruction instructions. RoguePrompt was developed and evaluated under a black-box threat model, with only API or user-interface access to the hosted models, and was tested on 313 real-world, hard-rejected prompts. Success was measured in terms of moderation bypass, instruction reconstruction, and execution when the relevant stage exceeded its automated criterion. RoguePrompt achieved average rates of 93.93% for filter bypass, 79.02% for reconstruction, and 70.18% for execution. These results demonstrate the effectiveness of layered prompt encoding while providing stage-level evidence of where multistage jailbreaks fail during moderation bypass, instruction reconstruction, and execution.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
- When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
- Deceiving Google's Perspective API Built for Detecting Toxic Comments
- Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks
- Prompt Injection attack against LLM-integrated Applications
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
- Training language models to follow instructions with human feedback
- A StrongREJECT for Empty Jailbreaks
- jina-embeddings-v3: Multilingual Embeddings With Task LoRA
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- Goal-guided Generative Prompt Injection Attack on Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs