From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation
cs.AI
Submitted: 2026-08-24
Updated: 2026-08-27
Comments: EMNLP 2026 Main Conference
Code: https://github.com/PangSMPang/NIS-Agent
License: http://creativecommons.org/licenses/by/4.0/
The gist: Web search agents powered by Large Language Models (LLMs) show strong promise, but deep research tasks expose a recurring failure mode: once an agent has produced a query, plan, or intermediate
Terminology
Abstract
Web search agents powered by Large Language Models (LLMs) show strong promise, but deep research tasks expose a recurring failure mode: once an agent has produced a query, plan, or intermediate conclusion, it becomes less objective when later judging the consequences of that same action. We term this phenomenon inertia bias. To make it measurable, we introduce the IBIS benchmark, which controls the search observations while varying whether the model is evaluating the outcome of its own prior action. We find that models are substantially worse when they "own" the preceding search step, showing that self-authored action history can systematically distort subsequent judgment. We further show that this bias propagates into two forms of system-level degradation: search noise at the worker level and contextual noise at the manager level. To address this problem, we propose NIS-Agent, which applies context isolation at the two decision points most vulnerable to inertia bias: webpage triage and final-answer validation. Across GAIA, WebWalkerQA, BrowseComp, and BrowseComp-zh, NIS-Agent achieves competitive performance while reducing token cost by 33% compared to our baseline. We further train an 8B model to be intrinsically more resistant to inertia bias; under the same NIS-Agent framework, it attains average performance comparable to GPT-4o on deep research benchmarks. Our code is publicly available at https://github.com/PangSMPang/NIS-Agent.
Sources
- Characterizing Deep Research: A Benchmark and Formal Definition
- Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization
- WebSailor: Navigating Super-human Reasoning for Web Agent
- WebThinker: Empowering Large Reasoning Models with Deep Research Capability
- WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
- APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
- Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- GAIA: a benchmark for General AI Assistants
- OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
- Large Language Models Cannot Self-Correct Reasoning Yet
- MemGPT: Towards LLMs as Operating Systems
- BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair
- Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving
- Discovering Language Model Behaviors with Model-Written Evaluations
- WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
- MemoBrain: Executive Memory as an Agentic Brain for Reasoning
- Qwen3 Technical Report
- Unveiling Confirmation Bias in Chain-of-Thought Reasoning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection