CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning
cs.AI
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- Gym-Anything: Turn any Software into an Agent Environment
- ClawGym: A Scalable Framework for Building Effective Claw Agents
- Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
- GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training
- Scaling Agent Learning via Experience Synthesis
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
- Mind2Web: Towards a Generalist Agent for the Web
- Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- The Rollout Infrastructure Tax in Coding-Agent Reinforcement Learning
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
- SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
- Qwen-CUA: Native Computer Use for (almost) Everything
- WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale
- Orchard: An Open-Source Agentic Modeling Framework
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
- AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
- ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection