Conversational recommendation systems (CRSs) enable users to use natural language feedback to control their recommendations, overcoming many of the challenges of traditional recommendation systems. However, the practical adoption of CRSs remains limited due to a lack of rich and diverse conversational training data that pairs user utterances with recommendations. To address this problem, we introduce a new method to generate synthetic training data by transforming curated item collections, such as playlists or movie watch lists, into item-seeking conversations. First, we use a biased random walk to generate a sequence of slates, or sets of item recommendations; then, we use a language model to generate corresponding user utterances. We demonstrate our approach by generating a conversational music recommendation dataset with over one million conversations, which were found to be consistent with relevant recommendations by a crowdsourced evaluation. Using the synthetic data to train a CRS, we significantly outperform standard retrieval baselines in offline and online evaluations.
翻译:对话式推荐系统(CRSs)使用户能够通过自然语言反馈控制推荐,克服了传统推荐系统的诸多挑战。然而,由于缺乏丰富多样的对话训练数据(将用户话语与推荐结果配对),CRSs的实际应用仍受到限制。为解决这一问题,我们提出一种新方法,通过将策划好的项目集合(如播放列表或电影观看列表)转化为寻求项目的对话,生成合成训练数据。首先,使用有偏随机游走生成一系列推荐列表(即项目推荐集合);然后,利用语言模型生成相应的用户话语。我们生成包含超过一百万次对话的对话式音乐推荐数据集来演示该方法,众包评估表明这些对话与相关推荐具有一致性。使用合成数据训练CRS后,我们在离线和在线评估中均显著优于标准检索基线。