Consumer speech recognition systems do not work as well for many people with speech diferences, such as stuttering, relative to the rest of the general population. However, what is not clear is the degree to which these systems do not work, how they can be improved, or how much people want to use them. In this paper, we frst address these questions using results from a 61-person survey from people who stutter and fnd participants want to use speech recognition but are frequently cut of, misunderstood, or speech predictions do not represent intent. In a second study, where 91 people who stutter recorded voice assistant commands and dictation, we quantify how dysfuencies impede performance in a consumer-grade speech recognition system. Through three technical investigations, we demonstrate how many common errors can be prevented, resulting in a system that cuts utterances of 79.1% less often and improves word error rate from 25.4% to 9.9%.
翻译:消费级语音识别系统对许多有言语差异(如口吃)的人群而言,效果不如普通人群。然而,尚不明确这些系统的失效程度、改进方式以及用户的使用意愿。本文首先通过一项针对61位口吃者的问卷调查回应上述问题,发现参与者希望使用语音识别系统,但常被中断、误解,或语音预测无法反映其意图。在第二项研究中,91位口吃者录制了语音助手指令与听写内容,我们量化了言语不流畅如何影响消费级语音识别系统的性能。通过三项技术研究,我们展示了许多常见错误可以被避免,最终使系统对语音片段的截断频率降低79.1%,词错误率从25.4%降至9.9%。