Generating realistic audio effects for movies and other media is a challenging task that is accomplished today primarily through physical techniques known as Foley art. Foley artists create sounds with common objects (e.g., boxing gloves, broken glass) in time with video as it is playing to generate captivating audio tracks. In this work, we aim to develop a deep-learning based framework that does much the same - observes video in it's natural sequence and generates realistic audio to accompany it. Notably, we have reason to believe this is achievable due to advancements in realistic audio generation techniques conditioned on other inputs (e.g., Wavenet conditioned on text). We explore several different model architectures to accomplish this task that process both previously-generated audio and video context. These include deep-fusion CNN, dilated Wavenet CNN with visual context, and transformer-based architectures. We find that the transformer-based architecture yields the most promising results, matching low-frequencies to visual patterns effectively, but failing to generate more nuanced waveforms.
翻译:为电影及其他媒体生成逼真音效是一项具有挑战性的任务,目前主要通过被称为拟音艺术的物理技术来完成。拟音艺术家利用常见物品(例如拳击手套、碎玻璃)随视频播放同步发声,从而创作出引人入胜的音轨。在本研究中,我们旨在开发一个基于深度学习的框架,实现类似功能——即观察视频的自然时序,并生成与之匹配的逼真音频。值得注意的是,我们有理由相信这一目标是可实现的,因为基于其他输入条件(例如,以文本为条件的Wavenet)的逼真音频生成技术已取得进展。我们探索了多种不同的模型架构来完成此任务,这些架构能够同时处理先前生成的音频和视频上下文,包括深度融合CNN、带视觉上下文的扩张Wavenet CNN,以及基于Transformer的架构。我们发现,基于Transformer的架构取得了最有前景的结果,能够有效地将低频信息与视觉模式相匹配,但尚无法生成更细微的波形。