Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses
cs.AI, cs.IR
Submitted: 2026-09-16
Updated: 2026-09-16
Code: https://github.com/mahsaama/AgenticSearchLens
License: http://creativecommons.org/licenses/by/4.0/
The gist: Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood.
Terminology
Abstract
Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-world user interactions (invivo) with controlled experiments using the same platform's models by their APIs (invitro). We investigate the quality of agentic decisions to invoke Web search, their strategies to formulate queries, the potential domain preferences in the search results they receive, and the choices they make when transforming search results into grounded responses. We find that Web-search decisions vary substantially across platforms and models, while more frequent Web-search invocation does not necessarily yield better response quality. We further show that conversational agents employ different complex querying strategies and that platform specific search engines return search results from their preferred domains. Finally, although responses are largely grounded in search results, some claims rely on uncited search results, raising concerns about attribution and reliability. Our findings have important implications for the design of future AI agents and Web search tools optimized for conversational retrieval.
Sources
- Deep Research Bench: Evaluating AI Web Research Agents
- How You Ask Matters! Adaptive RAG Robustness to Query Variations
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs
- Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and YouTube
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
- VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking
- WebGPT: Browser-assisted question-answering with human feedback
- Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents
- A Picture of Agentic Search
- Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
- R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling
- InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
- ShareChat: A Dataset of Chatbot Conversations in the Wild
- WildChat: 1M ChatGPT Interaction Logs in the Wild
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection