Most recent work in visual sound source localization relies on semantic audio-visual representations learned in a self-supervised manner, and by design excludes temporal information present in videos. While it proves to be effective for widely used benchmark datasets, the method falls short for challenging scenarios like urban traffic. This work introduces temporal context into the state-of-the-art methods for sound source localization in urban scenes using optical flow as a means to encode motion information. An analysis of the strengths and weaknesses of our methods helps us better understand the problem of visual sound source localization and sheds light on open challenges for audio-visual scene understanding.
翻译:近期的视觉声源定位研究主要依赖于以自监督方式学习的语义级视听表征,并在设计上排除了视频中存在的时序信息。尽管该方法在广泛使用的基准数据集上表现有效,但在城市交通等具有挑战性的场景中仍存在不足。本研究通过引入光流作为运动信息编码手段,将时序上下文融入面向城市场景的声源定位前沿方法。我们对方法优缺点的分析有助于更深入理解视觉声源定位问题,并为视听场景理解中的开放挑战提供启示。