WebArxiv: A Reproducible Benchmark for Evaluating Multimodal Web Agents on arXiv Tasks
cs.IR, cs.AI, cs.DB
Submitted: 2025-07-01
Updated: 2026-09-24
Comments: 14 pages, 5 figures, 7 tables
Code: https://github.com/Torantulino/Auto-GPT
License: http://creativecommons.org/licenses/by/4.0/
The gist: Foundation models now enable autonomous agents to interact with real-world websites, but existing benchmarks emphasize general-purpose browsing, underrepresent research-oriented environments and
Terminology
Abstract
Foundation models now enable autonomous agents to interact with real-world websites, but existing benchmarks emphasize general-purpose browsing, underrepresent research-oriented environments and scholarly discovery workflows, and often depend on live sites whose changing content and structure undermine reproducibility. arXiv provides a realistic, reproducible, hierarchically structured, information-centric testbed without privacy-sensitive interactions. We introduce WebArxiv, a static-snapshot benchmark comprising 510 time-invariant tasks, each with a unique deterministic ground truth. Its diverse, realistic scholarly tasks go beyond simple information lookup and rule following to emphasize multi-constraint paper retrieval, fine-grained content extraction, and cross-paper comparison. Evaluations of a range of foundation-model-based web agents show that WebArxiv remains challenging. Behavioral analysis reveals that agents over-rely on fixed interaction histories, causing incomplete or repetitive reasoning. We therefore equip agents with a lightweight dynamic-memory mechanism for adaptive retrieval and reasoning over relevant context. The benchmark and code are available at https://anonymous.4open.science/r/74E4423BVNW/README.md.
Sources
- Mind2Web: Towards a Generalist Agent for the Web
- Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
- REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites
- Gemini: A Family of Highly Capable Multimodal Models
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
- SeDi-Instruct: Enhancing Alignment of Language Models through Self-Directed Instruction Generation
- GFM-RAG: Graph Foundation Model for Retrieval Augmented Generation
- GPT-4 Technical Report
- WebCanvas: Benchmarking Web Agents in Online Environments
- ReAct: Synergizing Reasoning and Acting in Language Models
- Survey on Evaluation of LLM-based Agents
- AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
- LiteWebAgent: The Open-Source Suite for VLM-Based Web-Agent Applications
- GPT-4V(ision) is a Generalist Web Agent, if Grounded
- From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- There is no polynomial formula for the catenary and the tame degree of finitely generated monoids
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG