Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs
cs.CL, cs.AI
Submitted: 2025-11-24
Updated: 2026-08-28
Terminology
Sources
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- A Dataset for Answering Time-Sensitive Questions
- Dated Data: Tracing Knowledge Cutoffs in Large Language Models
- DailyQA: A Benchmark to Evaluate Web Retrieval Augmented LLMs Based on Capturing Real-World Changes
- The BrowserGym Ecosystem for Web Agent Research
- Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
- KnowRL: Teaching Language Models to Know What They Know
- Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
- Search Arena: Analyzing Search-Augmented LLMs
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- WeKnow-RAG: An Adaptive Approach for Retrieval-Augmented Generation Integrating Web Search and Knowledge Graphs
- Agent-SafetyBench: Evaluating the Safety of LLM Agents
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
- BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering