Descriptor generation methods using latent representations of encoder$-$decoder (ED) models with SMILES as input are useful because of the continuity of descriptor and restorability to the structure. However, it is not clear how the structure is recognized in the learning progress of ED models. In this work, we created ED models of various learning progress and investigated the relationship between structural information and learning progress. We showed that compound substructures were learned early in ED models by monitoring the accuracy of downstream tasks and input$-$output substructure similarity using substructure$-$based descriptors, which suggests that existing evaluation methods based on the accuracy of downstream tasks may not be sensitive enough to evaluate the performance of ED models with SMILES as descriptor generation methods. On the other hand, we showed that structure restoration was time$-$consuming, and in particular, insufficient learning led to the estimation of a larger structure than the actual one. It can be inferred that determining the endpoint of the structure is a difficult task for the model. To our knowledge, this is the first study to link the learning progress of SMILES by ED model to chemical structures for a wide range of chemicals.
翻译:利用编码器-解码器(ED)模型将SMILES作为输入生成的潜在表示进行描述符生成的方法,因其描述符的连续性和结构可恢复性而具有实用价值。然而,在ED模型的学习过程中,结构是如何被识别的尚不清楚。本研究构建了不同学习进度的ED模型,并探讨了结构信息与学习进度之间的关系。通过监测下游任务的准确率以及基于子结构描述符的输入-输出子结构相似性,我们发现化合物子结构在ED模型早期即被学习,这表明现有基于下游任务准确率的评估方法可能不足以敏感地评估以SMILES作为描述符生成方法的ED模型性能。另一方面,我们证明结构恢复是耗时的,特别是学习不足会导致估计的结构比实际结构更大。可以推断,确定结构终点对模型而言是一项困难任务。据我们所知,这是首个将ED模型对SMILES的学习进度与广泛化学品的化学结构相关联的研究。