In the absence of sensitive race and ethnicity data, researchers, regulators, and firms alike turn to proxies. In this paper, I train a Bidirectional Long Short-Term Memory (BiLSTM) model on a novel dataset of voter registration data from all 50 US states and create an ensemble that achieves up to 36.8% higher out of sample (OOS) F1 scores than the best performing machine learning models in the literature. Additionally, I construct the most comprehensive database of first and surname distributions in the US in order to improve the coverage and accuracy of Bayesian Improved Surname Geocoding (BISG) and Bayesian Improved Firstname Surname Geocoding (BIFSG). Finally, I provide the first high-quality benchmark dataset in order to fairly compare existing models and aid future model developers.
翻译:由于缺乏敏感的种族和民族数据,研究人员、监管机构和企业纷纷求助于代理变量。在本文中,我利用来自美国所有50个州的选民登记数据构建了一个新颖的数据集,并训练了一个双向长短期记忆(BiLSTM)模型,创建了一个集成模型,其在样本外(OOS)F1分数上比文献中性能最佳的机器学习模型高出36.8%。此外,我构建了美国最全面的名字和姓氏分布数据库,以提高贝叶斯改进姓氏地理编码(BISG)和贝叶斯改进名字-姓氏地理编码(BIFSG)的覆盖率和准确性。最后,我提供了首个高质量的基准数据集,以公平比较现有模型并帮助未来的模型开发者。