Constructing Challenging Browser-Use Tasks by Controlled Environment Interventions
cs.CL, cs.AI
Submitted: 2026-09-20
Updated: 2026-09-20
Code: https://github.com/browser-use/browser-use
Terminology
Sources
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
- ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress Conditions
- Language Models (Mostly) Know What They Know
- Kimi K2: Open Agentic Intelligence
- Kimi K2.5: Visual Agentic Intelligence
- ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
- Qwen3-VL Technical Report
- Stalled, Biased, and Confused: Uncovering Reasoning Failures in LLMs for Cloud-Based Root Cause Analysis
- From Grounding to Planning: Benchmarking Bottlenecks in Web Agents
- WIPI: A New Web Threat for LLM-Driven Web Agents
- The Rise and Potential of Large Language Model Based Agents: A Survey
- An Illusion of Progress? Assessing the Current State of Web Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering