We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline. In the proposed pipeline, multiple large language model (LLM) agents are created under various roles with specialized prompts and access to different parts of information, and the chat data is acquired by logging the conversation between the Listener LLM and the Recsys LLM. To cover various conversation scenarios, for each conversation, the Listener LLM is conditioned on a finetuned conversation goal. Finally, all the LLMs are multimodal with audio and images, allowing a simulation of multimodal recommendation and conversation. In the LLM-as-a-judge and subjective evaluation experiments, TalkPlayData 2 achieved the proposed goal in various aspects related to training a generative recommendation model for music. TalkPlayData 2 and its generation code are released at https://talkpl-ai.github.io.
翻译:我们提出了TalkPlayData 2,这是一个由可代理数据管道生成的多模态对话式音乐推荐合成数据集。在该管道中,多个大语言模型(LLM)代理被赋予不同角色,配有专用提示词并访问不同信息片段,通过记录Listener LLM与Recsys LLM之间的对话来获取聊天数据。为覆盖多种对话场景,每次对话中的Listener LLM均以微调后的对话目标为条件。最后,所有LLM均为支持音频与图像的多模态模型,从而实现对多模态推荐与对话的模拟。在"LLM作为评判者"实验及主观评估实验中,TalkPlayData 2在训练音乐生成式推荐模型的多个相关维度上达到了预期目标。TalkPlayData 2及其生成代码发布于https://talkpl-ai.github.io。