Topic modeling is a well-established technique for exploring text corpora. Conventional topic models (e.g., LDA) represent topics as bags of words that often require "reading the tea leaves" to interpret; additionally, they offer users minimal semantic control over topics. To tackle these issues, we introduce TopicGPT, a prompt-based framework that uses large language models (LLMs) to uncover latent topics within a provided text collection. TopicGPT produces topics that align better with human categorizations compared to competing methods: for example, it achieves a harmonic mean purity of 0.74 against human-annotated Wikipedia topics compared to 0.64 for the strongest baseline. Its topics are also more interpretable, dispensing with ambiguous bags of words in favor of topics with natural language labels and associated free-form descriptions. Moreover, the framework is highly adaptable, allowing users to specify constraints and modify topics without the need for model retraining. TopicGPT can be further extended to hierarchical topical modeling, enabling users to explore topics at various levels of granularity. By streamlining access to high-quality and interpretable topics, TopicGPT represents a compelling, human-centered approach to topic modeling.
翻译:主题建模是探索文本语料库的一项成熟技术。传统主题模型(如LDA)将主题表示为词袋,常常需要"解读茶叶"才能理解其含义;此外,它们为用户提供的主题语义控制能力非常有限。为解决这些问题,我们提出TopicGPT——一种基于提示的框架,利用大型语言模型(LLMs)从给定的文本集合中揭示潜在主题。与现有方法相比,TopicGPT生成的主题更符合人类分类标准:例如,在人工标注的维基百科主题上,其调和平均纯度达到0.74,而最强基线方法仅为0.64。该模型的主题可解释性更强,以自然语言标签及相关的自由描述替代了含义模糊的词袋表示。此外,该框架具有高度适应性,允许用户在不重新训练模型的情况下指定约束条件并修改主题。TopicGPT还可进一步扩展为层次化主题建模,使用户能在不同粒度级别上探索主题。通过简化对高质量、可解释主题的获取,TopicGPT为人本主题建模提供了一种引人注目的解决方案。