Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications
cs.HC, cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/Satwikram/OLLA
Terminology
Sources
- Qwen3-VL Technical Report
- Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
- The BrowserGym Ecosystem for Web Agent Research
- OSUniverse: Benchmark for Multimodal GUI-navigation AI Agents
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
- ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models
- OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
- OpenAI GPT-5 System Card
- An Illusion of Progress? Assessing the Current State of Web Agents
- Capability Self-Assessment in Large Language Models
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- Fine-Tuning Language Models from Human Preferences
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support