Nowadays, non-privacy small-scale motion detection has attracted an increasing amount of research in remote sensing in speech recognition. These new modalities are employed to enhance and restore speech information from speakers of multiple types of data. In this paper, we propose a dataset contains 7.5 GHz Channel Impulse Response (CIR) data from ultra-wideband (UWB) radars, 77-GHz frequency modulated continuous wave (FMCW) data from millimetre wave (mmWave) radar, and laser data. Meanwhile, a depth camera is adopted to record the landmarks of the subject's lip and voice. Approximately 400 minutes of annotated speech profiles are provided, which are collected from 20 participants speaking 5 vowels, 15 words and 16 sentences. The dataset has been validated and has potential for the research of lip reading and multimodal speech recognition.
翻译:近年来,非隐私性的小规模运动检测在遥感语音识别领域引起了越来越多的研究关注。这些新模态被用于增强和恢复来自多种数据类型说话者的语音信息。本文提出一个包含以下数据的数据集:来自超宽带雷达的7.5 GHz信道冲激响应数据、来自毫米波雷达的77 GHz调频连续波数据,以及激光数据。同时,采用深度相机记录受试者嘴唇和声音的标记点。该数据集提供了约400分钟带标注的语音档案,采集自20名参与者分别朗读5个元音、15个单词和16个句子的语音数据。该数据集已通过验证,在唇读及多模态语音识别研究方面具有潜在应用价值。