BreakFun: Jailbreaking LLMs via Object Instantiation under Simulated Code Execution
cs.CR, cs.AI, cs.CL
Submitted: 2025-10-19
Updated: 2026-09-18
Comments: Accepted to AACL-IJCNLP 2026 Findings
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues
- Multilingual Jailbreak Challenges in Large Language Models
- A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
- A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
- DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
- DeepInception: Hypnotize Large Language Model to Be Jailbreaker
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- FlipAttack: Jailbreak LLMs via Flipping
- CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
- Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
- Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
- Low-Resource Languages Jailbreak GPT-4
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
- WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response
- PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs