Machine learning models can make basic errors that are easily hidden within vast amounts of data. Such errors often run counter to human intuition referred to as "common sense". We thereby seek to characterize common sense for data-driven models, and quantify the extent to which a model has learned common sense. We propose a framework that integrates logic-based methods with statistical inference to derive common sense rules from a model's training data without supervision. We further show how to adapt models at test-time to reduce common sense rule violations and produce more coherent predictions. We evaluate our framework on datasets and models for three different domains. It generates around 250 to 300k rules over these datasets, and uncovers 1.5k to 26k violations of those rules by state-of-the-art models for the respective datasets. Test-time adaptation reduces these violations by up to 38% without impacting overall model accuracy.
翻译:机器学习模型可能犯下基本错误,这些错误往往隐匿在海量数据中,且常与人类直觉(即“常识”)相悖。为此,我们试图为数据驱动模型定义常识概念,并量化模型掌握常识的程度。我们提出一个框架,将基于逻辑的方法与统计推断相结合,无需监督即可从模型的训练数据中推导出常识规则。此外,我们展示了如何在测试时调整模型,以减少对常识规则的违反,从而生成更一致的预测结果。我们在三个不同领域的数据集和模型上评估了该框架。该框架在这些数据集上生成了约25万至30万条规则,并发现了针对各自数据集的现有最优模型违反这些规则的情况达1500至26000例。测试时调整可在不影响模型整体准确率的情况下,将违反率最多降低38%。