Traditional quantitative content analysis approach (human coding method) has weaknesses, such as assuming all human coders are equally accurate once the intercoder reliability for training reaches a threshold score. We applied the Biased-Annotator Competence Estimation (BACE) model (Tyler, 2021), which draws on Bayesian modeling to improve human coding. An important contribution of this model is it takes each coder's potential biases and reliability into consideration and treats the "true" label of each message as a latent parameter, with quantifiable estimation uncertainties. In contrast, in conventional human coding, each message will receive a fixed label without estimates for measurement uncertainties. In this extended abstract, we first summarize the weaknesses of conventional human coding; and then apply the BACE model to COVID-19 vaccine Twitter data and compare BACE with other statistical models; finally, we discuss how the BACE model can be applied to improve human coding of latent message features.
翻译:传统的定量内容分析方法(人工编码方法)存在缺陷,例如一旦训练达到交互编码者信度阈值,便假设所有人工编码者具有同等准确性。我们应用了基于贝叶斯建模改进人工编码的偏向标注者能力估计(BACE)模型(Tyler, 2021)。该模型的重要贡献在于,它将每位编码者的潜在偏向和可靠性纳入考量,并将每条信息的"真实"标签视为可量化估计不确定性的潜在参数。相比之下,传统人工编码中每条信息仅获得固定标签,无法估计测量误差。在本扩展摘要中,我们首先总结传统人工编码的局限性;继而将BACE模型应用于COVID-19疫苗推特数据,并将其与其他统计模型进行比较;最后讨论如何应用BACE模型改进对潜在信息特征的人工编码。