Supervised classification recognizes patterns in the data to separate classes of behaviours. Canonical solutions contain misclassification errors that are intrinsic to the numerical approximating nature of machine learning. The data analyst may minimize the classification error on a class at the expense of increasing the error of the other classes. The error control of such a design phase is often done in a heuristic manner. In this context, it is key to develop theoretical foundations capable of providing probabilistic certifications to the obtained classifiers. In this perspective, we introduce the concept of probabilistic safety region to describe a subset of the input space in which the number of misclassified instances is probabilistically controlled. The notion of scalable classifiers is then exploited to link the tuning of machine learning with error control. Several tests corroborate the approach. They are provided through synthetic data in order to highlight all the steps involved, as well as through a smart mobility application.
翻译:监督分类通过识别数据中的模式来区分行为类别。经典解中包含机器学习数值逼近本质所固有的误分类误差。数据分析师可能通过增加其他类别的误差来最小化某一类别的分类误差。此类设计阶段的误差控制通常采用启发式方法。在此背景下,发展能够为所得分类器提供概率保证的理论基础至关重要。基于此视角,我们提出概率安全区域的概念,用于描述输入空间中误分类实例数量受概率控制的子集。进而利用可扩展分类器的概念,将机器学习参数调整与误差控制联系起来。多项测试验证了该方法的有效性:通过合成数据阐明所有涉及步骤,并结合智能移动应用案例进行实证分析。