Sign languages are used as a primary language by approximately 70 million D/deaf people world-wide. However, most communication technologies operate in spoken and written languages, creating inequities in access. To help tackle this problem, we release ASL Citizen, the largest Isolated Sign Language Recognition (ISLR) dataset to date, collected with consent and containing 83,912 videos for 2,731 distinct signs filmed by 52 signers in a variety of environments. We propose that this dataset be used for sign language dictionary retrieval for American Sign Language (ASL), where a user demonstrates a sign to their own webcam with the aim of retrieving matching signs from a dictionary. We show that training supervised machine learning classifiers with our dataset greatly advances the state-of-the-art on metrics relevant for dictionary retrieval, achieving, for instance, 62% accuracy and a recall-at-10 of 90%, evaluated entirely on videos of users who are not present in the training or validation sets. An accessible PDF of this article is available at https://aashakadesai.github.io/research/ASL_Dataset__arxiv_.pdf
翻译:全世界约有7000万聋人将手语作为主要语言使用。然而,大多数通信技术基于口语和书面语言运作,造成了无障碍访问方面的不平等。为助力解决这一问题,我们发布了ASL Citizen数据集,这是迄今为止规模最大的孤立手语识别(ISLR)数据集。该数据集经用户知情同意后收集,包含由52位手语者在多种环境下拍摄的83,912个视频片段,涵盖2,731个不同手语词条。我们建议将该数据集用于美国手语(ASL)词典检索任务:用户通过自己的网络摄像头展示一个手语动作,系统从词典中检索匹配的手语词条。实验表明,利用本数据集训练监督式机器学习分类器,可在词典检索相关评价指标上显著提升当前最优性能——例如,在完全由未参与训练集或验证集的用户所拍摄的视频上,实现了62%的准确率和90%的召回率(Recall@10)。本文的无障碍PDF版本可通过https://aashakadesai.github.io/research/ASL_Dataset__arxiv_.pdf 获取。