Deep Research Agents Brings Deeper Harm
cs.CR, cs.CL
Submitted: 2025-10-13
Updated: 2026-09-05
Comments: COLM 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: We reveal that Deep Research (DR) agents systematically expose safety risks: simply submitting harmful queries that a standalone LLM would reject outright can elicit detailed and dangerous reports
Terminology
Abstract
We reveal that Deep Research (DR) agents systematically expose safety risks: simply submitting harmful queries that a standalone LLM would reject outright can elicit detailed and dangerous reports from DR agents. Empirical analysis reveals that the advantages that make DR agents powerful unintentionally make them vulnerable: both the research role assignment (e.g., Planner) and the multi-step execution mechanism weaken alignment of the base LLM, leading to severe safety breaches. Through linear probe analysis of internal hidden states, we demonstrate that role assignment suppresses refusal awareness by shifting representations away from safety boundaries. Besides, multi-step execution distributes harmfulness across individual steps, preventing alignment mechanisms from being activated throughout the research process. Exploiting these vulnerabilities, we design Intent Hijack (i.e., rephrasing harmful queries as academic research) and Plan Injection (i.e., manipulating execution plans) to further examine the safety risks of DR agents. Extensive experiments show that our methods achieve near-perfect compliance and elicit detailed, actionable reports that significantly exceed standalone LLM outputs in technical depth and applicability. These results demonstrate alarming misalignment in DR agents and underscore the urgent need for tailored alignment techniques. Our code is available in https://schen.app/deeper-harm.
Sources
- A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models
- Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
- AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
- Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models
- Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates
- Deep Research Agents: A Systematic Examination And Roadmap
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking
- Humanity's Last Exam
- Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks
- Multimodal Pragmatic Jailbreak on Text-to-image Models
- SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
- Search-o1: Agentic Search-Enhanced Large Reasoning Models
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Secrets of RLHF in Large Language Models Part I: PPO
- Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image
- Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models
- True Multimodal In-Context Learning Needs Attention to the Visual Context
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs