EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
cs.CL
Submitted: 2026-06-11
Updated: 2026-08-30
Code: https://github.com/networkx/networkx
Terminology
Sources
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
- WebSailor: Navigating Super-human Reasoning for Web Agent
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
- WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
- Tongyi DeepResearch Technical Report
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- Qwen3 Technical Report
- GLM-5: from Vibe Coding to Agentic Engineering
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering