Music recommendation for videos attracts growing interest in multi-modal research. However, existing systems focus primarily on content compatibility, often ignoring the users' preferences. Their inability to interact with users for further refinements or to provide explanations leads to a less satisfying experience. We address these issues with MuseChat, a first-of-its-kind dialogue-based recommendation system that personalizes music suggestions for videos. Our system consists of two key functionalities with associated modules: recommendation and reasoning. The recommendation module takes a video along with optional information including previous suggested music and user's preference as inputs and retrieves an appropriate music matching the context. The reasoning module, equipped with the power of Large Language Model (Vicuna-7B) and extended to multi-modal inputs, is able to provide reasonable explanation for the recommended music. To evaluate the effectiveness of MuseChat, we build a large-scale dataset, conversational music recommendation for videos, that simulates a two-turn interaction between a user and a recommender based on accurate music track information. Experiment results show that MuseChat achieves significant improvements over existing video-based music retrieval methods as well as offers strong interpretability and interactability.
翻译:视频音乐推荐在多模态研究中日益受到关注。然而,现有系统主要关注内容兼容性,往往忽略用户偏好。由于无法与用户交互以进行进一步优化或提供解释,导致用户体验欠佳。针对这些问题,我们提出MuseChat——首个基于对话的推荐系统,能够为视频提供个性化的音乐推荐。该系统包含两大核心功能及相应模块:推荐与推理。推荐模块以视频为输入,可选的附加信息包括先前推荐音乐及用户偏好,并检索出与上下文匹配的合适音乐。推理模块借助大语言模型(Vicuna-7B)的能力并扩展至多模态输入,能为推荐音乐提供合理的解释。为评估MuseChat的有效性,我们构建了大规模数据集——视频对话式音乐推荐,该数据集基于精确的音乐曲目信息模拟用户与推荐系统之间的两轮交互。实验结果表明,MuseChat较现有基于视频的音乐检索方法取得显著提升,同时具备强大的可解释性与交互性。