When training data is scarce, the incorporation of additional prior knowledge can assist the learning process. While it is common to initialize neural networks with weights that have been pre-trained on other large data sets, pre-training on more concise forms of knowledge has rather been overlooked. In this paper, we propose a novel informed machine learning approach and suggest to pre-train on prior knowledge. Formal knowledge representations, e.g. graphs or equations, are first transformed into a small and condensed data set of knowledge prototypes. We show that informed pre-training on such knowledge prototypes (i) speeds up the learning processes, (ii) improves generalization capabilities in the regime where not enough training data is available, and (iii) increases model robustness. Analyzing which parts of the model are affected most by the prototypes reveals that improvements come from deeper layers that typically represent high-level features. This confirms that informed pre-training can indeed transfer semantic knowledge. This is a novel effect, which shows that knowledge-based pre-training has additional and complementary strengths to existing approaches.
翻译:当培训数据稀少时,吸收更多的先前知识可以帮助学习过程。虽然在对其他大型数据集进行预先培训之前,通常会启动神经网络,其重量已经过其他大型数据集的预先培训,但关于更简洁的知识形式的预先培训却被忽略了。在本文件中,我们提出一种新的知情的机器学习方法,并建议对先前的知识进行预先培训。正式的知识表述,例如图表或方程式,首先可以转化为小型和精密的知识原型数据集。我们表明,关于这类知识原型的知情培训前培训(一) 加快学习过程,(二) 在没有足够的培训数据的情况下提高系统的一般化能力,以及(三) 提高模型的稳健性。分析模型中哪些部分受到原型的影响最大,表明改进来自通常代表高层次特征的更深层。这证实,知情的培训前展示确实可以转让语义知识。这是一个新效应,表明基于知识的培训前培训对现有方法具有额外和互补的长处。