This letter introduces an innovative method to enhance the quality of audio time stretching by precisely decomposing a sound into sines, transients, and noise and by improving the processing of the latter component. While there are established methods for time-stretching sines and transients with high quality, the manipulation of noise or residual components has lacked robust solutions in prior research. The proposed method combines sound decomposition with previous techniques for audio spectral resynthesis. The time-stretched noise component is achieved by morphing its time-interpolated spectral magnitude with a white-noise excitation signal. This method stands out for its simplicity, efficiency, and audio quality. The results of a subjective experiment affirm the superiority of this approach over current state-of-the-art methods across all evaluated stretch factors. The proposed technique notably excels in extreme stretching scenarios, signifying a substantial elevation in performance. The proposed method holds promise for a wide range of applications in slow-motion media content, such as music or sports video production.
翻译:本文提出一种创新方法,通过将声音精确分解为正弦分量、瞬态分量和噪声分量,并优化后者处理过程,来提升音频时间拉伸质量。尽管现有技术已能高质量实现正弦分量与瞬态分量的时间拉伸,但噪声或残余分量的处理在先前研究中缺乏稳健方案。本方法将声音分解技术与音频频谱再合成技术相结合,通过时间插值后的频谱幅度与白噪声激励信号进行变形处理,获得时间拉伸后的噪声分量。该方法在简洁性、计算效率与音质方面具有显著优势。主观实验结果表明,在所有测试拉伸比下,本方法均优于当前最先进技术。尤其在极端拉伸场景中,该方法展现出卓越性能提升,标志着技术水平的重大突破。本方法在慢动作媒体内容(如音乐或体育视频制作)领域具有广阔应用前景。