For tabular data sets, we explore data and model distillation, as well as data denoising. These techniques improve both gradient-boosting models and a specialized DNN architecture. While gradient boosting is known to outperform DNNs on tabular data, we close the gap for datasets with 100K+ rows and give DNNs an advantage on small data sets. We extend these results with input-data distillation and optimized ensembling to help DNN performance match or exceed that of gradient boosting. As a theoretical justification of our practical method, we prove its equivalence to classical cross-entropy knowledge distillation. We also qualitatively explain the superiority of DNN ensembles over XGBoost on small data sets. For an industry end-to-end real-time ML platform with 4M production inferences per second, we develop a model-training workflow based on data sampling that distills ensembles of models into a single gradient-boosting model favored for high-performance real-time inference, without performance loss. Empirical evaluation shows that the proposed combination of methods consistently improves model accuracy over prior best models across several production applications deployed worldwide.
翻译:对于表格数据集,我们探索了数据和模型蒸馏技术以及数据去噪方法。这些技术同时改进了梯度提升模型和专用深度神经网络架构。尽管梯度提升在表格数据上已知优于深度神经网络,但在具有10万行以上数据的数据集上我们缩小了这一差距,并使深度神经网络在小数据集上获得优势。我们通过输入数据蒸馏和优化集成进一步拓展这些结果,帮助深度神经网络性能匹配或超越梯度提升。作为我们实用方法的理论佐证,我们证明了其与经典交叉熵知识蒸馏的等价性。我们还从定性角度解释了为何深度神经网络集成在小数据集上优于XGBoost。针对每秒处理400万次生产推理的工业级端到端实时机器学习平台,我们开发了基于数据采样的模型训练工作流,该方法将模型集成知识蒸馏为单个梯度提升模型(该模型因高性能实时推理而受青睐)且不损失性能。实证评估表明,所提出的方法组合在全球部署的多个生产应用中持续提升了模型相对于先前最优模型的准确率。