In biomedical and public health association studies, binary outcome variables may be subject to misclassification, resulting in substantial bias in effect estimates. The feasibility of addressing binary outcome misclassification in regression models is often hindered by model identifiability issues. In this paper, we characterize the identifiability problems in this class of models as a specific case of "label switching" and leverage a pattern in the resulting parameter estimates to solve the permutation invariance of the complete data log-likelihood. Our proposed algorithm in binary outcome misclassification models does not require gold standard labels and relies only on the assumption that outcomes are correctly classified at least 50% of the time. A label switching correction is applied within estimation methods to recover unbiased effect estimates and to estimate misclassification rates in cases with one or more sequential observed outcomes. Open source software is provided to implement the proposed methods for single- and two-stage models. We give a detailed simulation study for our proposed methodology and apply these methods to data for single-stage modeling of the Medical Expenditure Panel Survey (MEPS) from 2020 and two-stage modeling of data from the Virginia Department of Criminal Justice Services.
翻译:在生物医学和公共卫生关联研究中,二分类结果变量可能受到错误分类的影响,从而导致效应估计出现显著偏倚。在回归模型中处理二分类结果错误分类的可行性常常受到模型可识别性问题的阻碍。本文将该类模型中的可识别性问题刻画为"标签交换"的一个特例,并利用参数估计结果中的某种模式来解决完全数据对数似然的置换不变性。我们提出的二分类结果错误分类算法不需要金标准标签,仅依赖于"结果被正确分类的概率至少为50%"这一假设。在估计方法中应用标签交换校正,以恢复无偏效应估计,并在存在一个或多个连续观测结果的情况下估计错误分类率。我们提供了开源的软件工具,以实现单阶段和两阶段模型的所提方法。对所提方法进行了详细的模拟研究,并将这些方法应用于以下数据:2020年医疗支出小组调查(MEPS)的单阶段建模,以及弗吉尼亚州刑事司法服务局数据的两阶段建模。