Scaling methods have long been utilized to simplify and cluster high-dimensional data. However, the general latent spaces across all predefined groups derived from these methods sometimes do not fall into researchers' interest regarding specific patterns within groups. To tackle this issue, we adopt an emerging analysis approach called contrastive learning. We contribute to this growing field by extending its ideas to multiple correspondence analysis (MCA) in order to enable an analysis of data often encountered by social scientists -- containing binary, ordinal, and nominal variables. We demonstrate the utility of contrastive MCA (cMCA) by analyzing two different surveys of voters in the U.S. and U.K. Our results suggest that, first, cMCA can identify substantively important dimensions and divisions among subgroups that are overlooked by traditional methods; second, for other cases, cMCA can derive latent traits that emphasize subgroups seen moderately in those derived by traditional methods.
翻译:缩放方法长期被用于简化高维数据并进行聚类。然而,这些方法在所有预定义群体中生成的通用潜在空间,有时并不符合研究者对群体内部特定模式的关注。为解决这一问题,我们采用了一种新兴分析方法——对比学习。我们通过将其思想扩展到多重对应分析(MCA)中,从而实现对社会科学研究者常遇到的数据(包含二元、有序及名义变量)的分析,为该领域的发展做出了贡献。通过分析美国和英国选民的两项不同调查,我们展示了对比性MCA(cMCA)的实用性。结果表明:首先,cMCA能够识别出传统方法所忽视的具有实质重要性的维度及子群体间的分化;其次,在其他案例中,cMCA能够推导出强调子群体的潜在特征,而这些特征在传统方法推导的结果中仅体现为中等程度。