The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge
cs.CL, cs.AI
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/alexgill321/2
Project page: https://alexgill321.github.io/KNOWS-benchmark
Terminology
Sources
- Emergence WebVoyager: Toward Consistent and Transparent Evaluation of (Web) Agents in The Wild
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- PentestJudge: Judging Agent Behavior Against Operational Requirements
- PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
- Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
- ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
- WebCanvas: Benchmarking Web Agents in Online Environments
- Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
- WebArxiv: A Reproducible Benchmark for Evaluating Multimodal Web Agents on arXiv Tasks
- Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
- OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
- OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- The Adoption and Usage of AI Agents: Early Evidence from Perplexity
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering