Temporal data (such as news articles or Twitter feeds) often consists of a mixture of long-lasting trends and popular but short-lasting topics of interest. A truly successful topic modeling strategy should be able to detect both types of topics and clearly locate them in time. In this paper, we first show that nonnegative CANDECOMP/PARAFAC decomposition (NCPD) is able to discover topics of variable persistence automatically. Then, we propose sparseness-constrained NCPD (S-NCPD) and its online variant in order to actively control the length of the learned topics effectively and efficiently. Further, we propose quantitative ways to measure the topic length and demonstrate the ability of S-NCPD (as well as its online variant) to discover short and long-lasting temporal topics in a controlled manner in semi-synthetic and real-world data including news headlines. We also demonstrate that the online variant of S-NCPD reduces the reconstruction error more rapidly than S-NCPD.
翻译:时序数据(如新闻文章或Twitter动态)通常包含长期存在趋势和短暂流行话题的混合。真正成功的话题建模策略应能同时检测这两种类型的话题,并准确定位其时间分布。本文首先证明非负CANDECOMP/PARAFAC分解(NCPD)能够自动发现具有可变持续期的话题。随后,我们提出稀疏约束NCPD(S-NCPD)及其在线变体,以主动、高效地控制学习话题的持续时间。进一步,我们提出量化话题长度的度量方法,并通过半合成数据和真实世界数据(包括新闻标题)验证S-NCPD(及其在线变体)能够以可控方式发现短期和长期时序话题的能力。同时证明S-NCPD在线变体比S-NCPD能更快降低重构误差。