Machine learning (ML) provides powerful tools for predictive modeling. ML's popularity stems from the promise of sample-level prediction with applications across a variety of fields from physics and marketing to healthcare. However, if not properly implemented and evaluated, ML pipelines may contain leakage typically resulting in overoptimistic performance estimates and failure to generalize to new data. This can have severe negative financial and societal implications. Our aim is to expand understanding associated with causes leading to leakage when designing, implementing, and evaluating ML pipelines. Illustrated by concrete examples, we provide a comprehensive overview and discussion of various types of leakage that may arise in ML pipelines.
翻译:机器学习(ML)为预测建模提供了强大工具。ML的普及源于其样本级预测能力,这一能力可应用于从物理学、市场营销到医疗保健等多个领域。然而,若未正确实施与评估,ML流水线中可能包含数据泄露,通常会导致性能估计过度乐观,且无法泛化至新数据——这对金融和社会可能造成严重的负面影响。本文旨在深化对设计、实施与评估ML流水线时引发泄露原因的理解。通过具体实例,我们系统梳理并讨论了ML流水线中可能出现的各类泄露现象。