Existing melody harmonization models have made great progress in improving the quality of generated harmonies, but most of them ignored the emotions beneath the music. Meanwhile, the variability of harmonies generated by previous methods is insufficient. To solve these problems, we propose a novel LSTM-based Hierarchical Variational Auto-Encoder (LHVAE) to investigate the influence of emotional conditions on melody harmonization, while improving the quality of generated harmonies and capturing the abundant variability of chord progressions. Specifically, LHVAE incorporates latent variables and emotional conditions at different levels (piece- and bar-level) to model the global and local music properties. Additionally, we introduce an attention-based melody context vector at each step to better learn the correspondence between melodies and harmonies. Objective experimental results show that our proposed model outperforms other LSTM-based models. Through subjective evaluation, we conclude that only altering the types of chords hardly changes the overall emotion of the music. The qualitative analysis demonstrates the ability of our model to generate variable harmonies.
翻译:现有的旋律和声生成模型在提升生成和声质量方面取得了显著进展,但大多数模型忽略了音乐背后的情感因素。同时,先前方法生成的和声多样性不足。为解决这些问题,我们提出了一种新颖的基于LSTM的分层变分自编码器(LHVAE),旨在探究情感条件对旋律和声生成的影响,同时提升生成和声质量并捕捉和弦进行的丰富多样性。具体而言,LHVAE在不同层级(乐曲级与节拍级)引入潜在变量和情感条件,以建模音乐的全局与局部属性。此外,我们每一步引入基于注意力的旋律上下文向量,以更好地学习旋律与和声之间的对应关系。客观实验结果表明,我们提出的模型优于其他基于LSTM的模型。通过主观评估,我们得出结论:仅改变和弦类型难以显著改变音乐的整体情感。定性分析证明我们的模型具备生成多样化和声的能力。