The availability of both structured and unstructured databases, such as electronic health data, social media data, patent data, and surveys that are often updated in real time, among others, has grown rapidly over the past decade. With this expansion, the statistical and methodological questions around data integration, or rather merging multiple data sources, has also grown. Specifically, the science of the ``data cleaning pipeline'' contains four stages that allow an analyst to perform downstream tasks, predictive analyses, or statistical analyses on ``cleaned data.'' This article provides a review of this emerging field, introducing technical terminology and commonly used methods.
翻译:过去十年间,结构化与非结构化数据库(包括电子健康数据、社交媒体数据、专利数据及常实时更新的调查数据等)的可用性快速增长。伴随这种扩张,围绕数据整合(即多数据源合并)的统计与方法论问题亦随之增多。具体而言,"数据清洗流程"这一学科包含四个阶段,使分析人员能够对"清洗后数据"执行下游任务、预测分析或统计分析。本文对此新兴领域进行综述,介绍技术术语与常用方法。