In social, medical, and behavioral research we often encounter datasets with a multilevel structure and multiple correlated dependent variables. These data are frequently collected from a study population that distinguishes several subpopulations with different (i.e., heterogeneous) effects of an intervention. Despite the frequent occurrence of such data, methods to analyze them are less common and researchers often resort to either ignoring the multilevel and/or heterogeneous structure, analyzing only a single dependent variable, or a combination of these. These analysis strategies are suboptimal: Ignoring multilevel structures inflates Type I error rates, while neglecting the multivariate or heterogeneous structure masks detailed insights. To analyze such data comprehensively, the current paper presents a novel Bayesian multilevel multivariate logistic regression model. The clustered structure of multilevel data is taken into account, such that posterior inferences can be made with accurate error rates. Further, the model shares information between different subpopulations in the estimation of average and conditional average multivariate treatment effects. To facilitate interpretation, multivariate logistic regression parameters are transformed to posterior success probabilities and differences between them. A numerical evaluation compared our framework to less comprehensive alternatives and highlighted the need to model the multilevel structure: Treatment comparisons based on the multilevel model had targeted Type I error rates, while single-level alternatives resulted in inflated Type I errors. A re-analysis of the Third International Stroke Trial data illustrated how incorporating a multilevel structure, assessing treatment heterogeneity, and combining dependent variables contributed to an in-depth understanding of treatment effects.
翻译:在社会、医学和行为科学研究中,我们常遇到具有多层结构和多个相关因变量的数据集。这些数据通常来自一个区分若干亚群(即处理效应存在异质性)的研究人群。尽管此类数据频繁出现,但针对其的分析方法相对较少,研究者常采用忽略多层和/或异质性结构、仅分析单个因变量,或结合上述策略的简化方法。这些分析策略并非最优:忽略多层结构会膨胀I类错误率,而忽视多元或异质性结构则会掩盖精细洞察。为全面分析此类数据,本文提出一种新颖的贝叶斯多层多元逻辑回归模型。该模型考虑了多层数据的聚类结构,从而能以准确的错误率进行后验推断。此外,模型在估计平均和条件平均多元处理效应时,在不同亚群间共享信息。为便于解释,多元逻辑回归参数被转换为后验成功概率及其差异。数值评估将本框架与简化替代方案进行对比,凸显了建模多层结构的必要性:基于多层模型的处理比较具有目标I类错误率,而单层替代方案导致I类错误膨胀。对第三次国际卒中试验数据的重新分析展示了纳入多层结构、评估处理异质性及整合因变量如何促进对处理效应的深入理解。