Users in consumption domains, like music, are often able to more efficiently provide preferences over a set of items (e.g. a playlist or radio) than over single items (e.g. songs). Unfortunately, this is an underexplored area of research, with most existing recommendation systems limited to understanding preferences over single items. Curating an item set exponentiates the search space that recommender systems must consider (all subsets of items!): this motivates conversational approaches-where users explicitly state or refine their preferences and systems elicit preferences in natural language-as an efficient way to understand user needs. We call this task conversational item set curation and present a novel data collection methodology that efficiently collects realistic preferences about item sets in a conversational setting by observing both item-level and set-level feedback. We apply this methodology to music recommendation to build the Conversational Playlist Curation Dataset (CPCD), where we show that it leads raters to express preferences that would not be otherwise expressed. Finally, we propose a wide range of conversational retrieval models as baselines for this task and evaluate them on the dataset.
翻译:在音乐等消费领域中,用户往往能更高效地表达对项目集(如歌单或电台)而非单个项目(如歌曲)的偏好。然而,这一研究领域尚未得到充分探索,现有推荐系统大多局限于理解用户对单个项目的偏好。整理项目集会指数级扩大推荐系统需考虑的搜索空间(所有项目的子集!),这催生了对话式方法——用户通过自然语言明确陈述或调整偏好,系统则通过自然语言引导偏好获取——作为理解用户需求的高效途径。我们将此任务定义为"对话式项目集整理",并提出一种新颖的数据收集方法,通过同时观测项目级和集合级反馈,在对话式场景中高效采集关于项目集的真实偏好。我们将该方法应用于音乐推荐领域,构建了"对话歌单整理数据集"(CPCD),实验表明该方法能引导标注者表达出在其他条件下不会表达的偏好。最后,我们提出多种对话式检索模型作为该任务的基线,并在该数据集上对其进行了评估。