Do Web Agents Investigate Before They Decide?
cs.AI
Submitted: 2026-02-05
Updated: 2026-09-08
Comments: 39 pages, 9 figures, 9 tables. Preprint
Code: https://github.com/user/chartgen
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Autonomous web agents are increasingly deployed in moderation and policy enforcement, where correct decisions often depend on evidence that is not immediately visible and must be actively
Terminology
Abstract
Autonomous web agents are increasingly deployed in moderation and policy enforcement, where correct decisions often depend on evidence that is not immediately visible and must be actively investigated. Yet existing benchmarks largely assume task critical information is immediately accessible. They do not measure investigative competence: recognizing when visible context is insufficient, retrieving hidden evidence, and integrating it into a final decision. We introduce MIRAGE, a benchmark of 750 multi step decision tasks across three domains: Wikipedia Forensics, Shopping Admin adjudication, and Reddit Moderation. Each task has two layers: a visible surface context that often points to the wrong action, and a hidden context, reachable only by active investigation, that contains the decisive evidence. We decompose agent performance into Investigation, Reasoning, and Decision Accuracy, complemented by an Investigative Hallucination Rate. We evaluate eight LLM agents across two model generations. Three patterns emerge. First, agents reach relevant pages but rarely extract the decisive evidence on them. Second, procedural hints improve investigation but do not consistently improve decisions on Wikipedia tasks, where decisive evidence often contradicts surface impressions. Third, 12.6% of trajectories cite fabricated facts. We call these the Navigation Discovery Gap, Collapse under Contradiction, and Investigative Hallucination. These patterns persist across model scale, generation, and reasoning architecture.
Sources
- Robust Hallucination Detection in LLMs via Adaptive Token Selection
- The Ouroboros of Benchmarking: Reasoning Evaluation in an Era of Saturation
- Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
- Enabling Large Language Models to Generate Text with Citations
- CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
- AgentBench: Evaluating LLMs as Agents
- CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
- Augmented Language Models: a Survey
- WebGPT: Browser-assisted question-answering with human feedback
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- Neural Operator Learning for Ultrasound Tomography Inversion
- WebCanvas: Benchmarking Web Agents in Online Environments
- WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation
- Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
- Wetting and Strain Engineering of 2D Materials on Nanopatterned Substrates
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- SafeArena: Evaluating the Safety of Autonomous Web Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection