SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents
cs.AI
Submitted: 2026-07-25
Updated: 2026-09-26
Code: https://github.com/ZJUSCL/SeekJudge
Terminology
Sources
- Gym-Anything: Turn any Software into an Agent Environment
- DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
- WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
- Mind2Web: Towards a Generalist Agent for the Web
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
- PRO-CUA: Process-Reward Optimization for Computer Use Agents
- Expanding Computation Spaces of LLMs at Inference Time
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
- On the Effects of Data Scale on UI Control Agents
- OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards
- Let's Verify Step by Step
- CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
- AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
- Autonomous Evaluation and Refinement of Digital Agents
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
- The Art of Building Verifiers for Computer Use Agents
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection