BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes
cs.CL
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/browser-use/browser-use
Terminology
Sources
- Qwen2.5-VL Technical Report
- GUICourse: From General Vision Language Models to Versatile GUI Agents
- UI-Venus Technical Report: Building High-performance UI Agents with RFT
- MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
- OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
- Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation
- OmniParser for Pure Vision Based GUI Agent
- UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
- GAIA: a benchmark for General AI Assistants
- NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild
- Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Qwen3-VL Technical Report
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
- OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
- A Survey on (M)LLM-Based GUI Agents
- ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- WebWorld: A Large-Scale World Model for Web Agent Training
- Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering