Developers and quality assurance testers often rely on manual testing to test accessibility features throughout the product lifecycle. Unfortunately, manual testing can be tedious, often has an overwhelming scope, and can be difficult to schedule amongst other development milestones. Recently, Large Language Models (LLMs) have been used for a variety of tasks including automation of UIs, however to our knowledge no one has yet explored their use in controlling assistive technologies for the purposes of supporting accessibility testing. In this paper, we explore the requirements of a natural language based accessibility testing workflow, starting with a formative study. From this we build a system that takes as input a manual accessibility test (e.g., ``Search for a show in VoiceOver'') and uses an LLM combined with pixel-based UI Understanding models to execute the test and produce a chaptered, navigable video. In each video, to help QA testers we apply heuristics to detect and flag accessibility issues (e.g., Text size not increasing with Large Text enabled, VoiceOver navigation loops). We evaluate this system through a 10 participant user study with accessibility QA professionals who indicated that the tool would be very useful in their current work and performed tests similarly to how they would manually test the features. The study also reveals insights for future work on using LLMs for accessibility testing.
翻译:开发者和质量保证测试人员通常依赖手动测试来验证产品生命周期中的功能无障碍性。然而,手动测试过程繁琐、覆盖范围过大,且难以与其他开发里程碑协调安排。近年来,大型语言模型已被广泛应用于用户界面自动化等任务,但据我们所知,尚无研究探索将其用于控制辅助技术以支持无障碍测试。本文通过一项形成性研究,探索了基于自然语言的无障碍测试工作流需求。基于此,我们构建了一个系统:输入手动无障碍测试指令(例如"在VoiceOver中搜索节目"),结合LLM与基于像素的UI理解模型执行测试,并生成带有章节标记的可导航视频。为辅助QA测试人员,我们在每个视频中应用启发式方法检测并标记无障碍问题(例如大文本启用时文本未缩放、VoiceOver导航循环)。通过10名无障碍QA专业人员参与的用户研究评估该系统,反馈表明该工具对其当前工作非常有用,且测试执行方式与手动测试相似。研究还揭示了未来利用LLM进行无障碍测试的洞察方向。