This paper highlights the importance of personalization in the current state of natural language understanding and generation and introduces the LaMP benchmark -- a novel benchmark for training and evaluating language models for producing personalized outputs. LaMP offers a comprehensive evaluation framework with diverse language tasks and multiple entries for each user profile. It consists of seven personalized tasks, spanning three classification and four text generation tasks. We also propose a retrieval augmentation approach that retrieves personalized items from user profiles to construct personalized prompts for large language models. Our baseline zero-shot and fine-tuned model results indicate that LMs utilizing profile augmentation outperform their counterparts that do not factor in profile information.
翻译:本文强调了个性化在当前自然语言理解与生成领域中的重要性,并介绍了LaMP基准——一个用于训练和评估语言模型以生成个性化输出的新型基准。LaMP提供了全面的评估框架,涵盖多种语言任务以及每个用户画像的多重条目。该基准包含七项个性化任务,涉及三项分类任务和四项文本生成任务。我们还提出了一种检索增强方法,从用户画像中检索个性化条目,以构建面向大语言模型的个性化提示。我们基于零样本和微调模型的基线结果表明,利用画像增强的语言模型在性能上优于未纳入画像信息的对应模型。