Recent studies on pronunciation scoring have explored the effect of introducing phone embeddings as reference pronunciation, but mostly in an implicit manner, i.e., addition or concatenation of reference phone embedding and actual pronunciation of the target phone as the phone-level pronunciation quality representation. In this paper, we propose to use linguistic-acoustic similarity to explicitly measure the deviation of non-native production from its native reference for pronunciation assessment. Specifically, the deviation is first estimated by the cosine similarity between reference phone embedding and corresponding acoustic embedding. Next, a phone-level Goodness of pronunciation (GOP) pre-training stage is introduced to guide this similarity-based learning for better initialization of the aforementioned two embeddings. Finally, a transformer-based hierarchical pronunciation scorer is used to map a sequence of phone embeddings, acoustic embeddings along with their similarity measures to predict the final utterance-level score. Experimental results on the non-native databases suggest that the proposed system significantly outperforms the baselines, where the acoustic and phone embeddings are simply added or concatenated. A further examination shows that the phone embeddings learned in the proposed approach are able to capture linguistic-acoustic attributes of native pronunciation as reference.
翻译:近期关于发音评分的研究探索了引入音素嵌入作为参考发音的效果,但大多采用隐式方式,即通过相加或拼接参考音素嵌入与目标音素的实际发音来生成音素级发音质量表征。本文提出利用语言-声学相似性显式度量非母语发音与母语参考之间的偏差,用于发音评估。具体而言,首先通过参考音素嵌入与对应声学嵌入之间的余弦相似性估算偏差值;其次引入音素级发音良好度(GOP)预训练阶段,引导基于相似性的学习过程以优化上述两类嵌入的初始化;最后采用基于Transformer的层级发音评分器,通过映射音素嵌入序列、声学嵌入序列及其相似性度量,预测最终的话语级评分。在非母语数据库上的实验结果表明,所提系统显著优于将声学嵌入与音素嵌入简单相加或拼接的基线系统。进一步分析显示,该方法习得的音素嵌入能够有效捕捉母语发音的语言-声学属性作为参考。