The prominence of a spoken word is the degree to which an average native listener perceives the word as salient or emphasized relative to its context. Speech prominence estimation is the process of assigning a numeric value to the prominence of each word in an utterance. These prominence labels are useful for linguistic analysis, as well as training automated systems to perform emphasis-controlled text-to-speech or emotion recognition. Manually annotating prominence is time-consuming and expensive, which motivates the development of automated methods for speech prominence estimation. However, developing such an automated system using machine-learning methods requires human-annotated training data. Using our system for acquiring such human annotations, we collect and open-source crowdsourced annotations of a portion of the LibriTTS dataset. We use these annotations as ground truth to train a neural speech prominence estimator that generalizes to unseen speakers, datasets, and speaking styles. We investigate design decisions for neural prominence estimation as well as how neural prominence estimation improves as a function of two key factors of annotation cost: dataset size and the number of annotations per utterance.
翻译:语音词语的显著性是指普通母语听者在语境中感知该词语突显或强调的程度。语音显著性估计是为语句中每个词语赋予一个数值的过程,用于表征其显著程度。这些显著性标注对语言分析以及训练自动化系统(如实现重点控制的文语转换或情感识别)具有重要价值。人工标注显著性耗时且昂贵,这推动了自动语音显著性估计方法的发展。然而,利用机器学习方法开发此类自动系统需要人工标注的训练数据。我们通过所构建的众包标注系统,收集并开源了LibriTTS数据集部分内容的众包标注结果。以这些标注为基准,我们训练了一个神经语音显著性估计器,该模型能够泛化至未见说话者、数据集及说话风格。我们研究了神经显著性估计的设计决策,并分析了神经显著性估计性能如何随标注成本的两个关键因素(数据集规模和每句话标注数量)的变化而提升。