Sounds reach one microphone in a stereo pair sooner than the other, resulting in an interaural time delay that conveys their directions. Estimating a sound's time delay requires finding correspondences between the signals recorded by each microphone. We propose to learn these correspondences through self-supervision, drawing on recent techniques from visual tracking. We adapt the contrastive random walk of Jabri et al. to learn a cycle-consistent representation from unlabeled stereo sounds, resulting in a model that performs on par with supervised methods on "in the wild" internet recordings. We also propose a multimodal contrastive learning model that solves a visually-guided localization task: estimating the time delay for a particular person in a multi-speaker mixture, given a visual representation of their face. Project site: https://ificl.github.io/stereocrw/
翻译:声音到达立体声对中某一麦克风的时间早于另一麦克风,由此产生的耳间时间延迟承载了声音的方向信息。估计声音的时间延迟需要找到各麦克风所记录信号之间的对应关系。我们借鉴近期视觉跟踪领域的先进技术,提出通过自监督学习来获取这些对应关系。我们改进Jabri等人提出的对比随机游走方法,从无标注的立体声信号中学习循环一致性表征,最终构建的模型在与真实互联网录音数据集上的有监督方法表现相当。此外,我们提出了一种多模态对比学习模型,可解决视觉引导定位任务:在多人混合语音中,通过给定人脸的视觉表征,估计特定人物的时延。项目网站:https://ificl.github.io/stereocrw/