Similar Accuracy, Unequal Evidence: Search APIs as Decision Surfaces for Tool-Using Agents
cs.CL
Submitted: 2026-07-11
Updated: 2026-09-13
License: http://creativecommons.org/licenses/by/4.0/
The gist: Search APIs expose ranked snippets, URLs, and metadata on which agents decide whether to answer, search again, or fetch pages.
Terminology
Abstract
Search APIs expose ranked snippets, URLs, and metadata on which agents decide whether to answer, search again, or fetch pages. We evaluate these interfaces as decision surfaces on a fixed sample of 100 questions from the 254-question SealQA-Hard subset, using one frozen GPT-5.4 agent, a fixed orchestration harness, and a shared page-fetch backend across Brave, Tavily, and Firecrawl. A Kimi-K2.6 oracle labels visible URL-level evidence; a separate answer audit yields 25, 25, and 26 correct answers out of 100. These counts indicate similar observed accuracy but do not establish equivalence. Under the tested configurations, Brave exposes more pre-fetch support alongside a larger snippet surface; Tavily has a larger rank-1 share among trajectory-pooled supporting observations; and Firecrawl is associated with broader exploration. First-search query-level metrics distinguish support availability from ranking, while contradiction exposure complements contradiction-to-gold ratios. A retrospective oracle union covers 44/100 questions, versus 26/100 for the best individual provider: 18 percentage points of headroom. The observed evidence and action differences motivate evaluating search APIs jointly with agent policy and retrieval budget.
Sources
- Evaluating Verifiability in Generative Search Engines
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- WebGPT: Browser-assisted question-answering with human feedback
- Automatic Evaluation of Attribution by Large Language Models
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Measuring and Narrowing the Compositionality Gap in Language Models
- ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
- FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- Over-Searching in Search-Augmented Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering