Dynamic topic modeling is widely used to analyze evolving trends in scientific literature, medical records, and social media. Traditional topic models represent each topic through a single probability vector on the multinomial simplex and implicitly couple word occurrence and repetition within one probabilistic mechanism. However, this formulation restricts the dependence structure among words and overlooks informative higher-order interactions, particularly in dynamic corpora with overlapping semantics. To address these limitations, we introduce a hypergraph representation of text where each document is modeled as a hyperedge connecting all co-occurring words, with repetition intensities encoded as node weights. This representation naturally separates word occurrence from repetition and induces a novel hypergraph-based multinomial distribution with a nonlinear normalization depending on the observed word set of each document. Building on this likelihood, we develop a dynamic topic modeling framework via structured low-rank factorizations with explicit temporal regularization on topic-word profiles. Moreover, we establish local convergence guarantees and derive non-asymptotic error bounds despite the intrinsic nonconvexity induced by bilinear factorization and document-specific nonlinear normalization. Numerical experiments on synthetic data and an application to the International Conference on Learning Representations (ICLR) corpus demonstrate consistent improvements over existing multinomial-based topic models.
翻译:动态主题建模广泛应用于科学文献、医疗记录和社交媒体中的演化趋势分析。传统主题模型通过多项单纯形上的单个概率向量表示每个主题,并在单一概率机制中隐含地耦合词语出现与重复。然而,这种表述限制了词语间的依赖结构,并忽略了具有信息价值的高阶交互——特别是在语义重叠的动态语料库中。针对这些局限性,我们引入文本的超图表示,将每个文档建模为连接所有共现词语的超边,并以节点权重编码重复强度。该表示自然地分离了词语出现与重复,并衍生出一种新型基于超图的多项分布,其非线性归一化取决于每个文档的观测词语集合。基于该似然函数,我们通过结构化低秩分解框架构建动态主题模型,并对主题-词谱进行显式时间正则化。进一步地,尽管双线性分解和文档特异性非线性归一化导致内在非凸性,我们仍建立了局部收敛保证并推导出非渐近误差界。合成数据的数值实验以及在《国际学习表征大会》(ICLR)语料库上的应用表明,该方法相较于现有基于多项式的主题模型具有持续改进效果。