In recent years, many recommender systems have utilized textual data for topic extraction to enhance interpretability. However, our findings reveal a noticeable deficiency in the coherence of keywords within topics, resulting in low explainability of the model. This paper introduces a novel approach called entropy regularization to address the issue, leading to more interpretable topics extracted from recommender systems, while ensuring that the performance of the primary task stays competitively strong. The effectiveness of the strategy is validated through experiments on a variation of the probabilistic matrix factorization model that utilizes textual data to extract item embeddings. The experiment results show a significant improvement in topic coherence, which is quantified by cosine similarity on word embeddings.
翻译:近年来,许多推荐系统利用文本数据进行主题提取以提升可解释性。然而,我们发现主题中关键词的连贯性存在显著不足,导致模型的可解释性较低。本文提出了一种称为熵正则化的新方法来解决这一问题,从而从推荐系统中提取更具可解释性的主题,同时确保主任务的性能保持竞争力。通过在一个利用文本数据提取项目嵌入的概率矩阵分解模型的变体上进行实验,验证了该策略的有效性。实验结果表明,主题连贯性显著提升,这一效果通过词嵌入的余弦相似度进行了量化。