Online health communities (OHCs) are vital for fostering peer support and improving health outcomes. Support groups within these platforms can provide more personalized and cohesive peer support, yet traditional support group formation methods face challenges related to scalability, static categorization, and insufficient personalization. To overcome these limitations, we propose two novel machine learning models for automated support group formation: the Group specific Dirichlet Multinomial Regression (gDMR) and the Group specific Structured Topic Model (gSTM). These models integrate user generated textual content, demographic profiles, and interaction data represented through node embeddings derived from user networks to systematically automate personalized, semantically coherent support group formation. We evaluate the models on a large scale dataset from MedHelp, comprising over 2 million user posts. Both models substantially outperform baseline methods including LDA, DMR, and STM in predictive accuracy (held out log likelihood), semantic coherence (UMass metric), and internal group consistency. The gDMR model yields group covariates that facilitate practical implementation by leveraging relational patterns from network structures and demographic data. In contrast, gSTM emphasizes sparsity constraints to generate more distinct and thematically specific groups. Qualitative analysis further validates the alignment between model generated groups and manually coded themes, showing the practical relevance of the models in informing groups that address diverse health concerns such as chronic illness management, diagnostic uncertainty, and mental health. By reducing reliance on manual curation, these frameworks provide scalable solutions that enhance peer interactions within OHCs, with implications for patient engagement, community resilience, and health outcomes.
翻译:在线健康社区(OHC)是促进同伴支持和改善健康结果的重要平台。其中的支持小组能够提供更个性化和凝聚力的同伴支持,但传统的支持小组形成方法面临可扩展性、静态分类和个性化不足的挑战。为克服这些局限,我们提出两种用于自动支持小组形成的新型机器学习模型:组特定狄利克雷多项式回归(gDMR)和组特定结构化主题模型(gSTM)。这些模型整合用户生成的文本内容、人口统计特征以及通过用户网络节点嵌入表示的交互数据,系统性地自动实现个性化、语义一致的支持小组形成。我们在包含超过200万条用户帖子的MedHelp大规模数据集上评估了模型性能。两种模型在预测准确性(保留对数似然)、语义连贯性(UMass指标)和组内一致性方面均显著优于LDA、DMR和STM等基线方法。gDMR模型通过利用网络结构中关系模式和人口统计数据生成组协变量,便于实际部署;而gSTM则通过稀疏性约束生成更鲜明且主题特异的组别。定性分析进一步验证了模型生成组与人工编码主题之间的对应关系,表明模型在识别涵盖慢性病管理、诊断不确定性和心理健康等多元健康议题的小组方面具有实际相关性。通过减少对人工标注的依赖,这些框架提供了可扩展的解决方案,能增强OHC中的同伴互动,对患者参与、社区韧性和健康结果具有重要启示。