Virtual sound synthesis is a technology that allows users to perceive spatial sound through headphones or earphones. However, accurate virtual sound requires an individual head-related transfer function (HRTF), which can be difficult to measure due to the need for a specialized environment. In this study, we proposed a method to generate HRTFs from one direction to the other. To this end, we used temporal convolutional neural networks (TCNs) to generate head-related impulse responses (HRIRs). To train the TCNs, publicly available datasets in the horizontal plane were used. Using the trained networks, we successfully generated HRIRs for directions other than the front direction in the dataset. We found that the proposed method successfully generated HRIRs for publicly available datasets. To test the generalization of the method, we measured the HRIRs of a new dataset and tested whether the trained networks could be used for this new dataset. Although the similarity evaluated by spectral distortion was slightly degraded, behavioral experiments with human participants showed that the generated HRIRs were equivalent to the measured ones. These results suggest that the proposed TCNs can be used to generate personalized HRIRs from one direction to another, which could contribute to the personalization of virtual sound.
翻译:虚拟声音合成技术允许用户通过耳机或耳塞感知空间声音。然而,精确的虚拟声音需要个体化的头相关传输函数(HRTF),由于需要特殊环境测量较为困难。本研究提出了一种从某一方向生成另一方向HRTF的方法。为此,我们使用时序卷积神经网络(TCNs)生成头相关脉冲响应(HRIRs)。训练TCNs时使用了水平面上的公开数据集。利用训练好的网络,我们成功生成了数据集中正面方向以外的HRIRs。研究发现,所提方法能成功为公开数据集生成HRIRs。为验证方法的泛化性,我们测量了新数据集的HRIRs并测试训练网络是否适用于该新数据集。尽管通过谱失真评估的相似度略有下降,但含人类参与者的行为实验表明,生成的HRIRs与实测结果等效。这些结果表明,所提出的TCNs可用于从某一方向生成另一方向的个性化HRIRs,这有助于实现虚拟声音的个性化。