This paper highlights the importance of personalization in large language models and introduces the LaMP benchmark -- a novel benchmark for training and evaluating language models for producing personalized outputs. LaMP offers a comprehensive evaluation framework with diverse language tasks and multiple entries for each user profile. It consists of seven personalized tasks, spanning three text classification and four text generation tasks. We additionally propose two retrieval augmentation approaches that retrieve personal items from each user profile for personalizing language model outputs. To this aim, we study various retrieval models, including term matching, semantic matching, and time-aware methods. Extensive experiments on LaMP for zero-shot and fine-tuned language models demonstrate the efficacy of the proposed retrieval augmentation approach and highlight the impact of personalization in various natural language tasks.
翻译:本文强调了个性化在大型语言模型中的重要性,并介绍了LaMP基准——一个用于训练和评估语言模型以生成个性化输出的新型基准。LaMP提供了一个全面的评估框架,包含多样的语言任务以及每个用户档案的多个条目。它由七个个性化任务组成,涵盖三个文本分类任务和四个文本生成任务。我们额外提出了两种检索增强方法,通过从每个用户档案中检索个性化条目来定制语言模型的输出。为此,我们研究了多种检索模型,包括词项匹配、语义匹配以及时间感知方法。在LaMP上对零样本和微调语言模型进行的大量实验,证明了所提检索增强方法的有效性,并突显了个性化在各种自然语言任务中的影响力。