We propose an unsupervised method to extract keywords and keyphrases from texts based on a pre-trained language model (LM) and Shannon's information maximization. Specifically, our method extracts phrases having the highest conditional entropy under the LM. The resulting set of keyphrases turns out to solve a relevant information-theoretic problem: if provided as side information, it leads to the expected minimal binary code length in compressing the text using the LM and an entropy encoder. Alternately, the resulting set is an approximation via a causal LM to the set of phrases that minimize the entropy of the text when conditioned upon it. Empirically, the method provides results comparable to the most commonly used methods in various keyphrase extraction benchmark challenges.
翻译:我们提出一种基于预训练语言模型(LM)和香农信息最大化的无监督方法,用于从文本中提取关键词和关键短语。具体而言,该方法提取在语言模型条件下具有最高条件熵的短语。所得到的关键短语集合可解决一个信息论相关问题:若将其作为侧信息提供,则在使用语言模型和熵编码器压缩文本时,能够实现预期的最小二进制码长度。该集合亦可视为通过因果语言模型对能使文本条件熵最小化的短语集合的近似。实验表明,该方法在多项关键短语提取基准挑战中的表现与最常用方法相当。