Strong but Brittle: Simple Attacks Subvert Reasoning-based Safety Guardrails
cs.CR, cs.CL
Submitted: 2025-10-13
Updated: 2026-09-05
Comments: OpenAI Red-teaming Challenge Winner and Oral Presentation | COLM 2026
Project page: https://chenxshuo.github.io/bag-of-tricks
License: http://creativecommons.org/licenses/by/4.0/
The gist: Open-weight Large Reasoning Models (LRMs) are approaching the capabilities of their frontier counterparts but pose significant safety concerns, as they are difficult to patch or monitor post-release.
Terminology
Abstract
Open-weight Large Reasoning Models (LRMs) are approaching the capabilities of their frontier counterparts but pose significant safety concerns, as they are difficult to patch or monitor post-release. To prevent misuse, reasoning-based safety guardrails, where models explicitly reason on safety justifications before answering, have become a promising primary defense, achieving near-perfect refusal rates on harmful queries. We show that this strong defense is alarmingly brittle and can be subverted by embarrassingly simple attacks to elicit extremely forbidden questions, such as `How to kill a man without being caught?' Specifically, we identify one systematic vulnerability from the reasoning-then-answer mechanism: the stage-transition logic that governs when safety reasoning begins and ends can be trivially manipulated. Based on this finding, we develop four simple yet effective red-teaming methods that systematically subvert different stages of the guardrails, either by bypassing reasoning entirely or exploiting it to produce targeted harmful content. These methods achieve attack success rates up to 90% across five benchmarks on multiple LRM families. Our findings reveal that reasoning-based guardrails are necessary but not sufficient and must be paired with robust triggering mechanisms and stronger base-model alignment. Code is in https://chenxshuo.github.io/bag-of-tricks/.
Sources
- Phi-4-reasoning Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
- Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
- Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
- Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
- OR-Bench: An Over-Refusal Benchmark for Large Language Models
- Practical Reasoning Interruption Attacks on Reasoning Large Language Models
- SAFER: Advancing Safety Alignment via Efficient Ex-Ante Reasoning
- A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models
- Deliberative Alignment: Reasoning Enables Safer Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- OpenAI o1 System Card
- Reasoning as an Adaptive Defense for Safety
- H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking
- Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
- AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
- Can an Individual Manipulate the Collective Decisions of Multi-Agents?
- Multimodal Pragmatic Jailbreak on Text-to-image Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs