The standard approach for neural topic modeling uses a variational autoencoder (VAE) framework that jointly minimizes the KL divergence between the estimated posterior and prior, in addition to the reconstruction loss. Since neural topic models are trained by recreating individual input documents, they do not explicitly capture the coherence between topic words on the corpus level. In this work, we propose a novel diversity-aware coherence loss that encourages the model to learn corpus-level coherence scores while maintaining a high diversity between topics. Experimental results on multiple datasets show that our method significantly improves the performance of neural topic models without requiring any pretraining or additional parameters.
翻译:神经主题建模的标准方法采用变分自编码器框架,该框架在重构损失的基础上联合最小化估计后验与先验之间的KL散度。由于神经主题模型通过重建单个输入文档进行训练,其并未显式捕获语料层面上主题词之间的一致性。本研究提出一种新颖的多样性感知一致性损失,在保持主题间高度多样性的同时,促使模型学习语料层面的一致性分数。多数据集实验结果表明,该方法无需任何预训练或额外参数即可显著提升神经主题模型的性能。