Automated assessment of speech intelligibility in hearing aid (HA) devices is of great importance. Our previous work introduced a non-intrusive multi-branched speech intelligibility prediction model called MBI-Net, which achieved top performance in the Clarity Prediction Challenge 2022. Based on the promising results of the MBI-Net model, we aim to further enhance its performance by leveraging Whisper embeddings to enrich acoustic features. In this study, we propose two improved models, namely MBI-Net+ and MBI-Net++. MBI-Net+ maintains the same model architecture as MBI-Net, but replaces self-supervised learning (SSL) speech embeddings with Whisper embeddings to deploy cross-domain features. On the other hand, MBI-Net++ further employs a more elaborate design, incorporating an auxiliary task to predict frame-level and utterance-level scores of the objective speech intelligibility metric HASPI (Hearing Aid Speech Perception Index) and multi-task learning. Experimental results confirm that both MBI-Net++ and MBI-Net+ achieve better prediction performance than MBI-Net in terms of multiple metrics, and MBI-Net++ is better than MBI-Net+.
翻译:助听器语音可懂度的自动评估具有重要意义。我们先前的工作提出了一种名为MBI-Net的非侵入式多分支语音可懂度预测模型,该模型在2022年Clarity预测挑战赛中取得了最佳性能。基于MBI-Net模型的优异结果,我们旨在通过利用Whisper嵌入来丰富声学特征,从而进一步提升其性能。在本研究中,我们提出了两种改进模型:MBI-Net+和MBI-Net++。MBI-Net+保持与MBI-Net相同的模型架构,但将自监督学习(SSL)语音嵌入替换为Whisper嵌入以部署跨域特征。另一方面,MBI-Net++采用了更精细的设计,引入辅助任务来预测客观语音可懂度指标HASPI(助听器语音感知指数)的帧级和话语级分数,并结合多任务学习。实验结果证实,MBI-Net++和MBI-Net+在多个指标上均比MBI-Net获得了更好的预测性能,且MBI-Net++优于MBI-Net+。