TimeWarp: Evaluating Web Agents by Revisiting the Past
cs.AI, cs.CL, cs.CV, cs.LG
Submitted: 2026-03-05
Updated: 2026-09-12
Code: https://github.com/web-arena-x/webarena
License: http://creativecommons.org/licenses/by/4.0/
The gist: As web agents close the gap with humans on benchmarks, one question arises: Do today's agents perform just as well on tomorrow's web? We introduce TimeWarp, a benchmark that emulates the evolving web.
Terminology
Abstract
As web agents close the gap with humans on benchmarks, one question arises: Do today's agents perform just as well on tomorrow's web? We introduce TimeWarp, a benchmark that emulates the evolving web. TimeWarp consists of three web environments, each with six UI versions spanning UI design, frontend code, and workflows from different eras of the internet. We pair TimeWarp with a set of complex, realistic tasks covering different forms of web navigation. Our experiments reveal that vision-based agents are vulnerable to changes, while text-based agents become brittle once fine-tuned on a single version. To address this, we propose TimeTraj, a new annotation method that uses plan distillation to collect trajectories across multiple versions. By training agents on teacher rollouts using our BC-variant, we achieve substantial performance gains: 20.4% to 37.7% for Qwen-3 4B and 0% to 27.0% for Llama-3.1 8B models. Our work helps study generalization across web designs and opens a new paradigm for collecting plans rather than trajectories to improve the robustness of web agents.
Sources
- The Llama 3 Herd of Models
- Training Deep Nets with Sublinear Memory Cost
- Pangu-Agent: A Fine-Tunable Generalist Agent with Structured Reasoning
- Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
- WebSuite: Systematically Evaluating Why Web Agents Fail
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
- Qwen2.5 Technical Report
- A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions
- Qwen3 Technical Report
- LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
- BEARCUBS: A benchmark for computer-using web agents
- Gemma 3 Technical Report
- InSTA: Towards Internet-Scale Training For Agents
- AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- An Illusion of Progress? Assessing the Current State of Web Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection