Binary classification is a fundamental task in machine learning, with applications spanning various scientific domains. Whether scientists are conducting fundamental research or refining practical applications, they typically assess and rank classification techniques based on performance metrics such as accuracy, sensitivity, and specificity. However, reported performance scores may not always serve as a reliable basis for research ranking. This can be attributed to undisclosed or unconventional practices related to cross-validation, typographical errors, and other factors. In a given experimental setup, with a specific number of positive and negative test items, most performance scores can assume specific, interrelated values. In this paper, we introduce numerical techniques to assess the consistency of reported performance scores and the assumed experimental setup. Importantly, the proposed approach does not rely on statistical inference but uses numerical methods to identify inconsistencies with certainty. Through three different applications related to medicine, we demonstrate how the proposed techniques can effectively detect inconsistencies, thereby safeguarding the integrity of research fields. To benefit the scientific community, we have made the consistency tests available in an open-source Python package.
翻译:二分类是机器学习中的基础任务,其应用涵盖多个科学领域。无论科研人员进行基础研究还是优化实际应用,通常都会基于准确率、灵敏度和特异度等性能指标来评估和排序分类技术。然而,报告的性能评分并非总能作为研究排名的可靠依据——这可能归因于与交叉验证相关的未公开或非常规操作、笔误及其他因素。在给定实验设置中,当正负测试样本数量确定时,大多数性能评分会呈现特定且相互关联的数值。本文提出一种数值技术,用于评估报告性能评分与假定实验设置的一致性。重要的是,该方法不依赖统计推断,而是通过数值方法确定性识别不一致性。通过三项医学领域的实际应用,我们展示了该技术如何有效检测不一致性,从而维护研究领域的完整性。为惠及科学共同体,我们已将一致性检验工具开源发布为Python程序包。