Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications
Satwik Ram Kodandaram, Monalika Padma Reddy, Xiaojun Bi, Jiawei Zhou, I. V. Ramakrishnan, Vikas Ashok
Abstract
Computer-use agents are emerging as a paradigm for agentic human-AI interaction, combining language reasoning with multi-modal interface grounding to operate GUIs. Yet their effectiveness for blind screen-reader users in real-world desktop workflows remains unclear. We present a three-week diary study with 8 blind users using OLLA, a screen-reader-accessible CUA prototype, collecting 1,258 commands across 12 applications with screenshots, UI trees, model responses, and action traces. We evaluate GPT-5 during deployment and re-execute the same commands with four additional models. GPT-5 achieved the highest success rate at 52.5%. Trace analysis reveals grounding, planning, constraint-tracking, and termination failures, while interviews reveal beyond-automation needs.
Create a lesson
Related papers
BuildOcc: A Large Language Model Occupant Agent Platform for Building Energy Research
Wooyoung Jung
Large Language Model-Driven Context-Aware Eco-Feedback Generation and Evaluation
Wooyoung Jung, Prosper Babon-Ayeng
The PIONEER Project: A PrIvacy companion for mOtivatioN and knowlEdge transfER
Simon Althaus, Nina Gerber, Sara Hahn et al.
Beyond Problem Solving: Large Language Models for Emotional and Reflective Support in Mathematics Learning
Vera Rief, Mirella Hladký, Minju Yoo et al.
EEG-based Visual Retrieval and Reconstruction: From Neurally Visible Optimal Layer to Hierarchical Diffusion Generation
Minyi Wang, Zhenqin Wu, Rihui Li
Decoding Decision Correctness from EEG Under High Cognitive Workload in Virtual Reality: Implications for Collaborative Brain-Computer Interface Teams
Christopher Baker, Stephen Hinton, Tom Reed et al.