ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices
cs.CR, cs.AI, cs.SE
Submitted: 2026-08-25
Updated: 2026-09-15
Code: https://github.com/facebookresearch/adepts-bench
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation
- Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions
- The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
- ADEPTS: A Capability Framework for Human-Centered Agent Design
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
- Qwen3-VL Technical Report
- AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
- DigiData: Training and Evaluating General-Purpose Mobile Control Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs