Nowadays, machine learning plays a key role in developing plenty of applications, e.g., smart homes, smart medical assistance, and autonomous driving. A major challenge of these applications is preserving high quality of the training and the serving data. Nevertheless, existing data cleaning methods cannot exploit context information. Thus, they usually fail to track shifts in the data distributions or the associated error profiles. To overcome these limitations, we introduce, in this paper, a novel method for automated tabular data cleaning powered by dynamic functional dependency rules extracted from a live context model. As a proof of concept, we create a smart home use case to collect data while preserving the context information. Using two different data sets, our evaluations show that the proposed cleaning method outperforms a set of baseline methods in terms of the detection and repair accuracy.
翻译:如今,机器学习在开发众多应用(如智能家居、智能医疗辅助及自动驾驶)中发挥着关键作用。这些应用面临的一大挑战在于保持训练数据与服务数据的高质量。然而,现有数据清洗方法无法利用上下文信息,因此通常难以追踪数据分布或关联错误模式的动态变化。为突破这些局限,本文提出一种新型自动化表格数据清洗方法,该方法借助从实时上下文模型中提取的动态函数依赖规则实现数据清洗。作为概念验证,我们构建了一个智能家居用例来收集数据并保留上下文信息。通过两组不同数据集的评估表明,所提出的清洗方法在检测与修复准确率上均优于多个基准方法。