Latent subgroup analysis is central to fields such as genomics, precision medicine, and social science, where the goal is to identify heterogeneous populations with distinct covariate structures or response behaviors. Mixture models provide a natural probabilistic framework for this task, representing the data-generating distribution as a weighted combination of subgroup-specific laws with unobserved labels. In high-dimensional regimes, these analyses face significant challenges. Often, only a small subset of covariates drives meaningful subgroup separation; the remaining variables may introduce noise or redundancy. Standard clustering methods typically treat all dimensions as equal, but in high-dimensional spaces, irrelevant coordinates can distort distances and obscure the low-dimensional structures defining latent classes. This paper introduces a Bayesian distilled clustering framework for high-dimensional mixture models. We propose that clustering should occur within a statistically justified subspace rather than the full ambient space. Our method utilizes a Bayesian variable selection model to estimate posterior inclusion probabilities, quantifying the evidence that each covariate contributes to subgroup separation or response behavior. A "distilled" covariate set is then identified by controlling the expected false-discovery proportion. Clustering is performed on this reduced subspace, followed by conditional independence diagnostics to examine subgroup-specific dependencies among selected variables. Critically, this framework is model-based: the distillation step is tied directly to the mixture structure and response model rather than a generic dimension-reduction criterion. This ensures the resulting subspace remains aligned with the scientific objective: identifying latent subgroups that differ in both distributional structure and behavior.


翻译:暂无翻译

0
下载
关闭预览

相关内容

专知会员服务
10+阅读 · 2021年10月1日
论文浅尝 | 一种嵌入效率极高的 node embedding 方式
开放知识图谱
13+阅读 · 2019年5月12日
Unsupervised Learning via Meta-Learning
CreateAMind
44+阅读 · 2019年1月3日
disentangled-representation-papers
CreateAMind
26+阅读 · 2018年9月12日
用 LDA 和 LSA 两种方法来降维和做 Topic 建模
AI研习社
13+阅读 · 2018年8月24日
量化投资与建模基于贝叶斯系列
量化投资与机器学习
15+阅读 · 2017年7月12日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
VIP会员
最新内容
受限仓库多智能体取送中的动态安全等待点选择
《国防技术管理》印度智库报告最新45页
专知会员服务
3+阅读 · 8月28日
《美陆军最新条令:保障行动》
专知会员服务
4+阅读 · 8月28日
算法战场:人工智能如何重新定义军事力量
专知会员服务
6+阅读 · 8月28日
《北约联邦式电子战云架构》
专知会员服务
6+阅读 · 8月27日
《美陆军野战手册:空域管理战术》
专知会员服务
10+阅读 · 8月27日
相关VIP内容
专知会员服务
10+阅读 · 2021年10月1日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员