GANStrument, exploiting GANs with a pitch-invariant feature extractor and instance conditioning technique, has shown remarkable capabilities in synthesizing realistic instrument sounds. To further improve the reconstruction ability and pitch accuracy to enhance the editability of user-provided sound, we propose HyperGANStrument, which introduces a pitch-invariant hypernetwork to modulate the weights of a pre-trained GANStrument generator, given a one-shot sound as input. The hypernetwork modulation provides feedback for the generator in the reconstruction of the input sound. In addition, we take advantage of an adversarial fine-tuning scheme for the hypernetwork to improve the reconstruction fidelity and generation diversity of the generator. Experimental results show that the proposed model not only enhances the generation capability of GANStrument but also significantly improves the editability of synthesized sounds. Audio examples are available at the online demo page.
翻译:GANStrument利用具有音高不变特征提取器和实例条件技术的GAN,在合成逼真的乐器声音方面展现了卓越能力。为进一步提升用户提供声音的重建能力与音高精度以增强可编辑性,我们提出了HyperGANStrument,该方法引入了一个音高不变超网络,以单次输入声音为条件,调节预训练GANStrument生成器的权重。超网络调制为生成器在输入声音重建过程中提供了反馈。此外,我们利用对抗性微调方案优化超网络,以提高生成器的重建保真度与生成多样性。实验结果表明,所提模型不仅增强了GANStrument的生成能力,还显著提升了合成声音的可编辑性。音频示例可在在线演示页面获取。