HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents
cs.CV
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/anthropics/claude-code
Terminology
Sources
- Qwen3-VL Technical Report
- RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
- Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
- ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
- UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents
- UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
- SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- OpenCUA: Open Foundations for Computer-Use Agents
- OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience
- MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models