Online health communities (OHCs) are vital for fostering peer support and improving health outcomes. Support groups within these platforms can provide more personalized and cohesive peer support, yet traditional support group formation methods face challenges related to scalability, static categorization, and insufficient personalization. To overcome these limitations, we propose two novel machine learning models for automated support group formation: the Group specific Dirichlet Multinomial Regression (gDMR) and the Group specific Structured Topic Model (gSTM). These models integrate user generated textual content, demographic profiles, and interaction data represented through node embeddings derived from user networks to systematically automate personalized, semantically coherent support group formation. We evaluate the models on a large scale dataset from MedHelp.org, comprising over 2 million user posts. Both models substantially outperform baseline methods including LDA, DMR, and STM in predictive accuracy (held out log likelihood), semantic coherence (UMass metric), and internal group consistency. The gDMR model yields group covariates that facilitate practical implementation by leveraging relational patterns from network structures and demographic data. In contrast, gSTM emphasizes sparsity constraints to generate more distinct and thematically specific groups. Qualitative analysis further validates the alignment between model generated groups and manually coded themes, showing the practical relevance of the models in informing groups that address diverse health concerns such as chronic illness management, diagnostic uncertainty, and mental health. By reducing reliance on manual curation, these frameworks provide scalable solutions that enhance peer interactions within OHCs, with implications for patient engagement, community resilience, and health outcomes.


翻译:在线健康社区(OHCs)对于促进同伴支持和改善健康结果至关重要。这些平台中的支持小组能够提供更个性化和凝聚力的同伴支持,然而传统的支持小组形成方法面临可扩展性、静态分类和个性化不足等挑战。为克服这些局限性,我们提出了两种用于自动形成支持小组的新型机器学习模型:组特定狄利克雷多项式回归(gDMR)和组特定结构化主题模型(gSTM)。这两种模型整合了用户生成的文本内容、人口统计资料以及通过用户网络节点嵌入表示的交互数据,以系统化地自动形成个性化、语义连贯的支持小组。我们在来自MedHelp.org的大规模数据集(包含超过200万条用户帖子)上评估了这些模型。在预测准确性(保留对数似然)、语义连贯性(UMass指标)和组内一致性方面,两种模型均显著优于LDA、DMR和STM等基线方法。gDMR模型通过利用网络结构和人口统计数据中的关系模式,生成了便于实际实施的组协变量。相比之下,gSTM模型强调稀疏性约束,以生成更独特和主题更明确的组。定性分析进一步验证了模型生成的组与人工编码主题之间的一致性,显示了这些模型在指导解决多种健康问题(如慢性病管理、诊断不确定性和心理健康)的小组方面的实际相关性。通过减少对人工管理的依赖,这些框架提供了可扩展的解决方案,增强了OHCs中的同伴互动,对患者参与、社区韧性和健康结果具有潜在影响。

0
下载
关闭预览

相关内容

【ICML2025】通过在线世界模型规划的持续强化学习
专知会员服务
20+阅读 · 2025年7月18日
常用的模型集成方法介绍:bagging、boosting 、stacking
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
博士论文 | 面向大模型推理的内存高效算法
专知会员服务
2+阅读 · 7月27日
美空军新型反无人机部队初探
专知会员服务
5+阅读 · 7月27日
《防空交战流程的概率建模研究》
专知会员服务
7+阅读 · 7月27日
ICML 2026 教程 | 数值优化理论还重要吗?
专知会员服务
6+阅读 · 7月26日
ICM 2026 | 陶哲轩:人工智能时代的数学
专知会员服务
9+阅读 · 7月26日
《反无人机交战场景下的战斗归零研究》
专知会员服务
7+阅读 · 7月26日
博士论文 | 用代码结构感知方法推进代码大模型
相关VIP内容
【ICML2025】通过在线世界模型规划的持续强化学习
专知会员服务
20+阅读 · 2025年7月18日
相关资讯
常用的模型集成方法介绍:bagging、boosting 、stacking
相关基金
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员