Developers and quality assurance testers often rely on manual testing to test accessibility features throughout the product lifecycle. Unfortunately, manual testing can be tedious, often has an overwhelming scope, and can be difficult to schedule amongst other development milestones. Recently, Large Language Models (LLMs) have been used for a variety of tasks including automation of UIs, however to our knowledge no one has yet explored their use in controlling assistive technologies for the purposes of supporting accessibility testing. In this paper, we explore the requirements of a natural language based accessibility testing workflow, starting with a formative study. From this we build a system that takes as input a manual accessibility test (e.g., ``Search for a show in VoiceOver'') and uses an LLM combined with pixel-based UI Understanding models to execute the test and produce a chaptered, navigable video. In each video, to help QA testers we apply heuristics to detect and flag accessibility issues (e.g., Text size not increasing with Large Text enabled, VoiceOver navigation loops). We evaluate this system through a 10 participant user study with accessibility QA professionals who indicated that the tool would be very useful in their current work and performed tests similarly to how they would manually test the features. The study also reveals insights for future work on using LLMs for accessibility testing.
翻译:开发人员和QA测试员通常依赖手动测试来验证整个产品生命周期中的无障碍功能。然而,手动测试过程繁琐、覆盖面广,且难以与其他开发里程碑协调安排。近年来,大型语言模型(LLMs)已被广泛应用于包括用户界面自动化在内的多种任务,但据我们所知,目前尚无人探索将其用于控制辅助技术以支持无障碍测试。本文通过一项形成性研究,探索了基于自然语言的无障碍测试工作流程的需求。基于此,我们构建了一个系统,该系统以手动无障碍测试描述(例如"在VoiceOver中搜索节目")为输入,结合LLM与基于像素的UI理解模型来执行测试,并生成带有章节划分的可导航视频。在每个视频中,为帮助QA测试员,我们应用启发式规则检测并标记无障碍问题(例如"大文本"功能开启时文字大小未增大、VoiceOver导航循环)。通过10名无障碍QA专业人员参与的用户研究评估表明,该工具在当前工作中极具实用性,且其执行测试的方式与手动测试高度相似。研究同时揭示了未来将LLM用于无障碍测试的洞见。