Fusing outputs from automatic speaker verification (ASV) and spoofing countermeasure (CM) is expected to make an integrated system robust to zero-effort imposters and synthesized spoofing attacks. Many score-level fusion methods have been proposed, but many remain heuristic. This paper revisits score-level fusion using tools from decision theory and presents three main findings. First, fusion by summing the ASV and CM scores can be interpreted on the basis of compositional data analysis, and score calibration before fusion is essential. Second, the interpretation leads to an improved fusion method that linearly combines the log-likelihood ratios of ASV and CM. However, as the third finding reveals, this linear combination is inferior to a non-linear one in making optimal decisions. The outcomes of these findings, namely, the score calibration before fusion, improved linear fusion, and better non-linear fusion, were found to be effective on the SASV challenge database.
翻译:融合自动说话人验证(ASV)与欺骗防御(CM)系统的输出,有望构建一个对零成本冒名者和合成欺骗攻击均具有鲁棒性的集成系统。目前已提出多种评分级融合方法,但多数仍具有启发性。本文借助决策理论工具重新审视评分级融合,并提出三项主要发现。首先,基于组合数据分析理论可解释ASV与CM评分求和的融合方式,且融合前的评分校准至关重要。其次,该理论解释推导出一种改进的融合方法,即对ASV与CM的对数似然比进行线性组合。然而,第三项发现表明,在实现最优决策方面,这种线性组合逊色于非线性组合。基于这些发现形成的技术方案——包括融合前评分校准、改进的线性融合及更优的非线性融合——在SASV挑战数据库中被验证具有显著效果。