Many apps have basic accessibility issues, like missing labels or low contrast. Automated tools can help app developers catch basic issues, but can be laborious or require writing dedicated tests. We propose a system, motivated by a collaborative process with accessibility stakeholders at a large technology company, to generate whole app accessibility reports by combining varied data collection methods (e.g., app crawling, manual recording) with an existing accessibility scanner. Many such scanners are based on single-screen scanning, and a key problem in whole app accessibility reporting is to effectively de-duplicate and summarize issues collected across an app. To this end, we developed a screen grouping model with 96.9% accuracy (88.8% F1-score) and UI element matching heuristics with 97% accuracy (98.2% F1-score). We combine these technologies in a system to report and summarize unique issues across an app, and enable a unique pixel-based ignore feature to help engineers and testers better manage reported issues across their app's lifetime. We conducted a qualitative evaluation with 18 accessibility-focused engineers and testers which showed this system can enhance their existing accessibility testing toolkit and address key limitations in current accessibility scanning tools.
翻译:许多应用存在基本的可达性问题,例如缺失标签或对比度不足。自动化工具虽能帮助开发者检测基本问题,但往往操作繁琐或需要编写专门的测试。我们提出一种系统,其设计源于与一家大型科技公司可达性相关利益方的协作流程,通过将多种数据收集方法(如应用爬取、手动录制)与现有可达性扫描器相结合,生成完整的应用可达性报告。许多此类扫描器基于单屏幕扫描,而全应用可达性报告的关键问题在于有效去重并汇总跨屏幕收集的问题。为此,我们开发了准确率达96.9%(F1分数为88.8%)的屏幕分组模型,以及准确率达97%(F1分数为98.2%)的UI元素匹配启发式算法。我们将这些技术整合到一个系统中,用于报告和汇总应用中的唯一问题,并引入独特的基于像素的忽略功能,帮助工程师和测试人员更好地管理应用生命周期中上报的问题。我们与18位专注于可达性的工程师和测试人员进行了定性评估,结果表明该系统能够增强现有的可达性测试工具集,并解决当前可达性扫描工具的关键局限性。