Sensitising language models (LMs) to external context helps them to more effectively capture the speaking patterns of individuals with specific characteristics or in particular environments. This work investigates to what extent rich character and film annotations can be leveraged to personalise LMs in a scalable manner. We then explore the use of such models in evaluating context specificity in machine translation. We build LMs which leverage rich contextual information to reduce perplexity by up to 6.5% compared to a non-contextual model, and generalise well to a scenario with no speaker-specific data, relying on combinations of demographic characteristics expressed via metadata. Our findings are consistent across two corpora, one of which (Cornell-rich) is also a contribution of this paper. We then use our personalised LMs to measure the co-occurrence of extra-textual context and translation hypotheses in a machine translation setting. Our results suggest that the degree to which professional translations in our domain are context-specific can be preserved to a better extent by a contextual machine translation model than a non-contextual model, which is also reflected in the contextual model's superior reference-based scores.
翻译:赋予语言模型(LMs)对外部语境的敏感性,有助于其更有效地捕捉具有特定特征或处于特定环境中的个体的说话模式。本研究探讨如何充分利用丰富的角色和影片注释,以可扩展的方式实现语言模型的个性化。进而,我们探索利用此类模型评估机器翻译中语境特异性的可能性。通过构建利用丰富语境信息的语言模型,我们的模型相比非语境模型可将困惑度降低高达6.5%,并在无说话者特定数据的场景下表现出良好的泛化能力——这依赖于通过元数据表示的人口统计学特征组合。该结论在两个语料库中具有一致性,其中Cornell-rich语料库亦是本文的贡献之一。随后,我们使用个性化语言模型,在机器翻译场景中测量文本外语境与翻译假设的共现性。结果表明,在该领域内,语境化机器翻译模型相比非语境模型能更好地保留专业翻译的语境特异性程度,这一现象亦反映在语境模型更优的基于参考的评分中。