Being able to express our thoughts, feelings, and ideas to one another is essential for human survival and development. A considerable portion of the population encounters communication obstacles in environments where hearing is the primary means of communication, leading to unfavorable effects on daily activities. An autonomous sign language recognition system that works effectively can significantly reduce this barrier. To address the issue, we proposed a large scale dataset namely Multi-View Bangla Sign Language dataset (MV- BSL) which consist of 115 glosses and 350 isolated words in 15 different categories. Furthermore, We have built a recurrent neural network (RNN) with attention based bidirectional gated recurrent units (Bi-GRU) architecture that models the temporal dynamics of the pose information of an individual communicating through sign language. Human pose information, which has proven effective in analyzing sign pattern as it ignores people's body appearance and environmental information while capturing the true movement information makes the proposed model simpler and faster with state-of-the-art accuracy.
翻译:能够将我们的思想、情感和想法相互表达,是人类生存与发展的关键。相当一部分人口在以听觉为主要交流方式的环境中面临沟通障碍,这对日常活动产生了不利影响。一套有效运作的自主手语识别系统能够显著减少这一障碍。为解决此问题,我们提出了一个大型数据集,即多视角孟加拉手语数据集(MV-BSL),该数据集包含15个不同类别的115个词汇和350个孤立词。此外,我们构建了一个基于注意力机制的双向门控循环单元(Bi-GRU)架构的循环神经网络(RNN),该网络对通过手语进行交流的个体姿态信息的时间动态进行建模。人类姿态信息在分析手语模式中已被证明有效,因为它忽略了个人的身体外观和环境信息,同时捕捉了真实的运动信息,这使得所提出的模型更简单、更快速,并达到了最先进的准确率。