In the field of chemical structure recognition, the task of converting molecular images into graph structures and SMILES string stands as a significant challenge, primarily due to the varied drawing styles and conventions prevalent in chemical literature. To bridge this gap, we proposed MolNexTR, a novel image-to-graph deep learning model that collaborates to fuse the strengths of ConvNext, a powerful Convolutional Neural Network variant, and Vision-TRansformer. This integration facilitates a more nuanced extraction of both local and global features from molecular images. MolNexTR can predict atoms and bonds simultaneously and understand their layout rules. It also excels at flexibly integrating symbolic chemistry principles to discern chirality and decipher abbreviated structures. We further incorporate a series of advanced algorithms, including improved data augmentation module, image contamination module, and a post-processing module to get the final SMILES output. These modules synergistically enhance the model's robustness against the diverse styles of molecular imagery found in real literature. In our test sets, MolNexTR has demonstrated superior performance, achieving an accuracy rate of 81-97%, marking a significant advancement in the domain of molecular structure recognition. Scientific contribution: MolNexTR is a novel image-to-graph model that incorporates a unique dual-stream encoder to extract complex molecular image features, and combines chemical rules to predict atoms and bonds while understanding atom and bond layout rules. In addition, it employs a series of novel augmentation algorithms to significantly enhance the robustness and performance of the model.
翻译:在化学结构识别领域,将分子图像转化为图结构和SMILES字符串是一项重大挑战,这主要源于化学文献中普遍存在的多样化绘制风格和惯例。为弥合这一差距,我们提出了MolNexTR——一种创新的图像到图深度学习模型,它协同融合了ConvNext(一种强大的卷积神经网络变体)与Vision-Transformer的优势。这种整合使得模型能够更精细地从分子图像中提取局部和全局特征。MolNexTR可同时预测原子和化学键,并理解它们的布局规则。它还能灵活整合符号化学原理,以识别手性并解析缩写结构。我们进一步引入了一系列先进算法,包括改进的数据增强模块、图像污染模块以及后处理模块,以获取最终的SMILES输出。这些模块协同增强了模型对真实文献中多样化分子图像风格的鲁棒性。在我们的测试集中,MolNexTR展现了优越性能,准确率达到81-97%,标志着分子结构识别领域的重大进展。科学贡献:MolNexTR是一种新颖的图像到图模型,它整合了独特的双流编码器用于提取复杂分子图像特征,并结合化学规则来预测原子和化学键,同时理解原子和化学键的布局规则。此外,它采用了一系列创新增强算法,显著提升了模型的鲁棒性和性能。