Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
cs.AI
Submitted: 2026-07-09
Updated: 2026-09-07
Comments: Accepted at ICML 2026. 9 pages, 6 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information.
Terminology
Abstract
Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information during an auditing process, we introduce overthinking: the process of using reasoning task vectors to amplify the propensity to think out loud of reasoning models. Given the parameters of a non-reasoning instruct model M and reasoning-distilled model R, we define the overthinking model as θ O α = θ M + α(θ R - θ M), where α> 1 amplifies reasoning beyond the pure reasoning model R. Additionally, we introduce new layer-wise attenuation strategies that selectively amplify reasoning without losing quality and coherence of model outputs. We demonstrate that overthinking models are more likely to reveal hidden information across four experimental settings, across 2B-32B models. Our findings suggest that reasoning amplification may surface secrets or unintended behaviors acquired during training up to 10 times more frequently than the original reasoning model. How secrets surface depends on the secret type: some require perturbation along the reasoning direction, while others yield to any sufficiently large weight perturbation.
Sources
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- Steering Language Models with Weight Arithmetic
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
- Qwen3-VL Technical Report
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Reasoning Models Don't Always Say What They Think
- Are DeepSeek R1 And Other Reasoning Models More Faithful?
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals
- Eliciting Secret Knowledge from Language Models
- Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
- Noise Injection Systemically Degrades Large Language Model Safety Guardrails
- Model evaluation for extreme risks
- Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection