OverThink: Slowdown Attacks on Reasoning LLMs
cs.LG, cs.CR
Submitted: 2025-02-04
Updated: 2026-09-17
Code: https://github.com/akumar2709/OVERTHINK_public
License: http://creativecommons.org/licenses/by/4.0/
The gist: A reasoning language model (RLM) generates costly reasoning tokens, often hidden from the users, that help it excel at many tasks.
Terminology
Abstract
A reasoning language model (RLM) generates costly reasoning tokens, often hidden from the users, that help it excel at many tasks. Our Overthink attack targets RLM-based applications (such as chatbots or coding agents) that rely on external context by forcing these models to generate substantially more reasoning tokens while still producing contextually correct answers. An adversary conducts the attack by injecting decoy reasoning problems into available content, optimized to elicit a large number of tokens. We craft decoy challenges (using Markov decision processes, language translation, or graphic comprehension) that appear benign individually yet have an adversarial impact when inserted in the context, allowing them to easily evade safety filters. We evaluate Overthink on proprietary and open-source reasoning models across the FreshQA, SQuAD, and MuSR datasets, where we observe 13x, 46x, and 12x increases, respectively. We also explore multimodal attacks using images, which cause up to a 2.7x increase in reasoning, as well as attacks on coding agents by injecting decoys into skills, README files, and code, resulting in up to a 17x increase. We explore several defenses and evaluate their efficacy against different attack strategies, highlighting that defending against Overthink is nontrivial. Finally, we show that newer generations of RLMs, while showing a drastic increase in per-token cost, also exhibit up to a 2.3x increase in reasoning tokens, leaving them more vulnerable to Overthink attacks.
Sources
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Chinchilla Scaling: A replication attempt
- From Persona to Personalization: A Survey on Role-Playing Language Agents
- SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
- Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
- Gradient-based Jailbreak Images for Multimodal Fusion Models
- Antidistillation Sampling
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
- Excessive Reasoning Attack on Reasoning LLMs
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- A Survey of Reasoning with Foundation Models
- Kimi K2: Open Agentic Intelligence
- Hype, Sustainability, and the Price of the Bigger-is-Better Paradigm in AI
- Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
- OckBench: Measuring the Efficiency of LLM Reasoning
- Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
- Denial-of-Service Poisoning Attacks against Large Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- The Curious Case of Neural Text Degeneration
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks