CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering
cs.AI, cs.SE
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/THUDM/slime
Terminology
Sources
- Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
- Programming with Pixels: Can Computer-Use Agents do Software Engineering?
- Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
- Evaluating Large Language Models Trained on Code
- GameDevBench: Evaluating Agentic Capabilities Through Game Development
- Mind2Web: Towards a Generalist Agent for the Web
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
- On the Effects of Data Scale on UI Control Agents
- WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
- GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- OpenCUA: Open Foundations for Computer-Use Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection