In recent years, dynamic parameterization of acoustic environments has raised increasing attention in the field of audio processing. One of the key parameters that characterize the local room acoustics in isolation from orientation and directivity of sources and receivers is the geometric room volume. Convolutional neural networks (CNNs) have been widely selected as the main models for conducting blind room acoustic parameter estimation, which aims to learn a direct mapping from audio spectrograms to corresponding labels. With the recent trend of self-attention mechanisms, this paper introduces a purely attention-based model to blindly estimate room volumes based on single-channel noisy speech signals. We demonstrate the feasibility of eliminating the reliance on CNN for this task and the proposed Transformer architecture takes Gammatone magnitude spectral coefficients and phase spectrograms as inputs. To enhance the model performance given the task-specific dataset, cross-modality transfer learning is also applied. Experimental results demonstrate that the proposed model outperforms traditional CNN models across a wide range of real-world acoustics spaces, especially with the help of the dedicated pretraining and data augmentation schemes.
翻译:近年来,声学环境的动态参数化在音频处理领域日益受到关注。表征局部房间声学特性(与声源及接收器的方向性和指向性无关)的关键参数之一为几何房间体积。卷积神经网络(CNN)被广泛选作实施盲房间声学参数估计的主要模型,其目标是从音频频谱图中学习到对应标签的直接映射。基于自注意力机制的最新发展趋势,本文提出一种纯注意力模型,利用单通道含噪语音信号盲估计房间体积。我们论证了在此任务中消除对CNN依赖的可行性,所提出的Transformer架构以伽马通幅度谱系数和相位频谱图作为输入。为增强模型在任务特定数据集上的性能,还应用了跨模态迁移学习。实验结果表明,所提模型在真实声学空间广泛范围内优于传统CNN模型,尤其在专用预训练与数据增强方案的辅助下效果更佳。