Human-spoken questions are critical to evaluating the performance of spoken question answering (SQA) systems that serve several real-world use cases including digital assistants. We present a new large-scale community-shared SQA dataset, HeySQuAD that consists of 76k human-spoken questions and 97k machine-generated questions and corresponding textual answers derived from the SQuAD QA dataset. The goal of HeySQuAD is to measure the ability of machines to understand noisy spoken questions and answer the questions accurately. To this end, we run extensive benchmarks on the human-spoken and machine-generated questions to quantify the differences in noise from both sources and its subsequent impact on the model and answering accuracy. Importantly, for the task of SQA, where we want to answer human-spoken questions, we observe that training using the transcribed human-spoken and original SQuAD questions leads to significant improvements (12.51%) over training using only the original SQuAD textual questions.
翻译:人类口语问题对于评估服务于包括数字助手在内的多种实际应用场景的口语问答(SQA)系统的性能至关重要。我们提出了一个全新的大规模社区共享SQA数据集——HeySQuAD,该数据集包含7.6万个由人类口述的问题和9.7万个机器生成的问题,以及从SQuAD问答数据集中提取的相应文本答案。HeySQuAD旨在衡量机器理解含噪声口语问题并准确回答这些问题的能力。为此,我们对人类口述问题和机器生成问题进行了广泛的基准测试,以量化两种来源的噪声差异及其对模型性能和回答准确性的后续影响。重要的是,在旨在回答人类口语问题的SQA任务中,我们观察到使用转录的人类口语问题和原始SQuAD问题进行训练,相比仅使用原始SQuAD文本问题训练,性能提升了12.51%。