The research in Deep Learning applications in sound and music computing have gathered an interest in the recent years; however, there is still a missing link between these new technologies and on how they can be incorporated into real-world artistic practices. In this work, we explore a well-known Deep Learning architecture called Variational Autoencoders (VAEs). These architectures have been used in many areas for generating latent spaces where data points are organized so that similar data points locate closer to each other. Previously, VAEs have been used for generating latent timbre spaces or latent spaces of symbolic music excepts. Applying VAE to audio features of timbre requires a vocoder to transform the timbre generated by the network to an audio signal, which is computationally expensive. In this work, we apply VAEs to raw audio data directly while bypassing audio feature extraction. This approach allows the practitioners to use any audio recording while giving flexibility and control over the aesthetics through dataset curation. The lower computation time in audio signal generation allows the raw audio approach to be incorporated into real-time applications. In this work, we propose three strategies to explore latent spaces of audio and timbre for sound design applications. By doing so, our aim is to initiate a conversation on artistic approaches and strategies to utilize latent audio spaces in sound and music practices.
翻译:近年来,深度学习方法在声音与音乐计算领域的研究引起了广泛关注。然而,这些新技术如何融入实际艺术实践仍存在认知断层。本研究探索了一种著名的深度学习架构——变分自编码器(VAEs)。该架构已广泛应用于生成潜在空间,其中数据点通过相似性聚类分布。现有研究多将VAE应用于音色潜在空间或符号音乐片段的潜在空间生成。若将VAE应用于音色音频特征,需借助声码器将网络生成的音色转换为音频信号,这会带来高昂的计算成本。本研究直接对原始音频数据应用VAE,绕过音频特征提取步骤。该方法允许实践者使用任意音频录音,并通过数据集策展实现对美学特征的灵活控制。音频信号生成的较低计算开销使得原始音频方法可集成至实时应用场景。本文针对声音设计应用提出了三种探索音频与音色潜在空间的策略,旨在开启关于利用潜在音频空间进行声音与音乐实践的艺术方法与策略的讨论。