We introduce Gull, a generative multifunctional audio codec. Gull is a general purpose neural audio compression and decompression model which can be applied to a wide range of tasks and applications such as real-time communication, audio super-resolution, and codec language models. The key components of Gull include (1) universal-sample-rate modeling via subband modeling schemes motivated by recent progress in audio source separation, (2) gain-shape representations motivated by traditional audio codecs, (3) improved residual vector quantization modules for simpler training, (4) elastic decoder network that enables user-defined model size and complexity during inference time, (5) built-in ability for audio super-resolution without the increase of bitrate. We compare Gull with existing traditional and neural audio codecs and show that Gull is able to achieve on par or better performance across various sample rates, bitrates and model complexities in both subjective and objective evaluation metrics.
翻译:我们提出Gull,一种生成式多功能音频编解码器。Gull是一个通用的神经音频压缩与解压缩模型,可应用于实时通信、音频超分辨率和编解码语言模型等多种任务及应用场景。Gull的关键组件包括:(1)基于近期音频源分离进展的子带建模方案实现通用采样率建模;(2)受传统音频编解码器启发的增益-形状表示;(3)改进的残差向量量化模块以简化训练过程;(4)弹性解码器网络,支持推理时用户自定义模型规模与复杂度;(5)内置无需增加比特率即可实现音频超分辨率的能力。我们将Gull与现有传统及神经音频编解码器进行对比,结果表明Gull在不同采样率、比特率和模型复杂度下,在主观与客观评价指标中均能达到或超越现有性能水平。