Machine learning approaches to modelling analog audio effects have seen intensive investigation in recent years, particularly in the context of non-linear time-invariant effects such as guitar amplifiers. For modulation effects such as phasers, however, new challenges emerge due to the presence of the low-frequency oscillator which controls the slowly time-varying nature of the effect. Existing approaches have either required foreknowledge of this control signal, or have been non-causal in implementation. This work presents a differentiable digital signal processing approach to modelling phaser effects in which the underlying control signal and time-varying spectral response of the effect are jointly learned. The proposed model processes audio in short frames to implement a time-varying filter in the frequency domain, with a transfer function based on typical analog phaser circuit topology. We show that the model can be trained to emulate an analog reference device, while retaining interpretable and adjustable parameters. The frame duration is an important hyper-parameter of the proposed model, so an investigation was carried out into its effect on model accuracy. The optimal frame length depends on both the rate and transient decay-time of the target effect, but the frame length can be altered at inference time without a significant change in accuracy.
翻译:近年来,机器学习方法在模拟音频效果建模领域得到了广泛研究,尤其针对吉他放大器等非线性时不变效果器。然而,对于移相器等调制效果器,低频振荡器控制着效果随时间缓慢变化的特性,这带来了新的挑战。现有方法要么需要预先知晓该控制信号,要么在实现上存在非因果性。本文提出一种可微分数字信号处理方法用于移相效果器建模,能够联合学习效果器的基础控制信号及其时变频谱响应。该模型通过短时音频帧处理,在频域实现时变滤波器,其传递函数基于典型模拟移相器电路拓扑结构。实验表明,该模型可训练用于模拟参考模拟设备,同时保留可解释且可调节的参数。帧时长是本模型的关键超参数,因此我们系统研究了其对模型精度的影响。最优帧长度取决于目标效果器的速率和瞬态衰减时间,但推理时调整帧长度不会显著改变模型精度。