In many multi-microphone algorithms, an estimate of the relative transfer functions (RTFs) of the desired speaker is required. Recently, a computationally efficient RTF vector estimation method was proposed for acoustic sensor networks, assuming that the spatial coherence (SC) of the noise component between a local microphone array and multiple external microphones is low. Aiming at optimizing the output signal-to-noise ratio (SNR), this method linearly combines multiple RTF vector estimates, where the complex-valued weights are computed using a generalized eigenvalue decomposition (GEVD). In this paper, we perform a theoretical bias analysis for the SC-based RTF vector estimation method with multiple external microphones. Assuming a certain model for the noise field, we derive an analytical expression for the weights, showing that the optimal model-based weights are real-valued and only depend on the input SNR in the external microphones. Simulations with real-world recordings show a good accordance of the GEVD-based and the model-based weights. Nevertheless, the results also indicate that in practice, estimation errors occur which the model-based weights cannot account for.
翻译:在许多多麦克风算法中,需要估计目标说话人的相对传递函数(RTF)向量。近年来,针对声传感器网络提出了一种计算高效的RTF向量估计方法,该方法假设局部麦克风阵列与多个外部麦克风之间噪声分量的空间相干性(SC)较低。该方法以优化输出信噪比(SNR)为目标,线性组合多个RTF向量估计值,其中复值权重通过广义特征值分解(GEVD)计算。本文针对多外部麦克风场景下基于SC的RTF向量估计方法进行理论偏差分析。在特定声场模型假设下,我们推导了权重的解析表达式,表明基于最优模型的权重为实值,且仅取决于外部麦克风的输入信噪比。基于真实场景录制的仿真实验表明,GEVD计算权重与模型推导权重具有良好一致性。然而,结果也显示在实际应用中会出现模型权重无法解释的估计误差。