The standard approach for neural topic modeling uses a variational autoencoder (VAE) framework that jointly minimizes the KL divergence between the estimated posterior and prior, in addition to the reconstruction loss. Since neural topic models are trained by recreating individual input documents, they do not explicitly capture the coherence between topic words on the corpus level. In this work, we propose a novel diversity-aware coherence loss that encourages the model to learn corpus-level coherence scores while maintaining a high diversity between topics. Experimental results on multiple datasets show that our method significantly improves the performance of neural topic models without requiring any pretraining or additional parameters.
翻译:针对神经主题建模,标准方法采用变分自编码器(VAE)框架,在最小化重建损失的同时,联合优化估计后验与先验之间的KL散度。由于神经主题模型通过重建单个输入文档进行训练,它们并未显式捕获语料层面主题词之间的连贯性。本研究提出了一种新颖的面向多样性的连贯性损失,该损失在保持主题间高度多样性的同时,激励模型学习语料层面的连贯性得分。多数据集上的实验结果表明,该方法无需预训练或额外参数,即可显著提升神经主题模型的性能。