Language has always been one of humanity's defining characteristics. Visual Language Identification (VLI) is a relatively new field of research that is complex and largely understudied. In this paper, we present a preliminary study in which we use linguistic information as a soft biometric trait to enhance the performance of a visual (auditory-free) identification system based on lip movement. We report a significant improvement in the identification performance of the proposed visual system as a result of the integration of these data using a score-based fusion strategy. Methods of Deep and Machine Learning are considered and evaluated. To the experimentation purposes, the dataset called laBial Articulation for the proBlem of the spokEn Language rEcognition (BABELE), consisting of eight different languages, has been created. It includes a collection of different features of which the spoken language represents the most relevant, while each sample is also manually labelled with gender and age of the subjects.
翻译:语言一直是人类最具定义性的特征之一。视觉语言识别(VLI)是一个相对较新的研究领域,具有复杂性且尚未被充分研究。本文提出了一项初步研究,利用语言信息作为软生物特征,以增强基于唇动的无声视觉识别系统的性能。通过基于得分的融合策略整合这些数据,我们报告了所提出的视觉系统在识别性能上的显著提升。研究中考虑并评估了深度学习和机器学习方法。为实验目的,我们创建了名为“唇部发音用于口语语言识别问题(BABELE)”的数据集,包含八种不同语言。该数据集收集了多种特征,其中口语语言是最相关的特征,同时每个样本还手动标注了受试者的性别和年龄。