Automatic speech recognition systems based on deep learning are mainly trained under empirical risk minimization (ERM). Since ERM utilizes the averaged performance on the data samples regardless of a group such as healthy or dysarthric speakers, ASR systems are unaware of the performance disparities across the groups. This results in biased ASR systems whose performance differences among groups are severe. In this study, we aim to improve the ASR system in terms of group robustness for dysarthric speakers. To achieve our goal, we present a novel approach, sample reweighting with sample affinity test (Re-SAT). Re-SAT systematically measures the debiasing helpfulness of the given data sample and then mitigates the bias by debiasing helpfulness-based sample reweighting. Experimental results demonstrate that Re-SAT contributes to improved ASR performance on dysarthric speech without performance degradation on healthy speech.
翻译:基于深度学习的自动语音识别系统主要采用经验风险最小化(ERM)进行训练。由于ERM仅关注数据样本的平均性能而忽略健康或构音障碍说话者等特定群体,ASR系统无法感知不同群体间的性能差异,导致产生群体间性能差异严重的偏差化ASR系统。本研究旨在提升ASR系统对构音障碍说话者的群体鲁棒性。为实现该目标,我们提出了一种创新方法——基于样本亲和性测试的样本重加权(Re-SAT)。该方法通过系统性地度量给定数据样本的消偏有效性,并据此进行基于消偏有效性的样本加权来缓解偏差。实验结果表明,Re-SAT在保持健康语音识别性能不下降的同时,有效提升了构音障碍语音的ASR性能。