Personalisation of language models for dialogue sensitises them to better capture the speaking patterns of people of specific characteristics, and/or in specific environments. However, rich character annotations are difficult to come by and to successfully leverage. In this work, we release and describe a novel set of manual annotations for 863 speakers from the popular Cornell Movie Dialog Corpus, including features like characteristic quotes and character descriptions, and a set of six automatically extracted metadata for over 95% of the featured films. We perform extensive experiments on two corpora and show that such annotations can be effectively used to personalise language models, reducing perplexity by up to 8.5%. Our method can be applied even zero-shot for speakers for whom no prior training data is available, by relying on combinations of characters' demographic characteristics. Since collecting such metadata is costly, we also contribute a cost-benefit analysis to highlight which annotations were most cost-effective relative to the reduction in perplexity.
翻译:对话语言模型的个性化使其能够更精确地捕捉具有特定特征和/或处于特定环境中人群的说话模式。然而,丰富的角色标注既难以获取又难以有效利用。本文中,我们发布并描述了一组针对康奈尔电影对话语料库中863名说话人的新型人工标注数据,包括角色经典台词和人物描述等特征,以及针对95%以上影片的六项自动提取元数据。我们在两个语料库上进行了广泛实验,证明此类标注可有效用于语言模型个性化,使困惑度降低高达8.5%。即便对于无先验训练数据的说话人,我们的方法也能基于角色人口统计特征的组合实现零样本应用。由于此类元数据收集成本高昂,我们还提供了成本效益分析,以明确哪些标注相对于困惑度降低具有最高成本效益。