Many contemporary generative models of molecules are variational auto-encoders of molecular graphs. One term in their training loss pertains to reconstructing the input, yet reconstruction capabilities of state-of-the-art models have not yet been thoroughly compared on a large and chemically diverse dataset. In this work, we show that when several state-of-the-art generative models are evaluated under the same conditions, their reconstruction accuracy is surprisingly low, worse than what was previously reported on seemingly harder datasets. However, we show that improving reconstruction does not directly lead to better sampling or optimization performance. Failed reconstructions from the MoLeR model are usually similar to the inputs, assembling the same motifs in a different way, and possess similar chemical properties such as solubility. Finally, we show that the input molecule and its failed reconstruction are usually mapped by the different encoders to statistically distinguishable posterior distributions, hinting that posterior collapse may not fully explain why VAEs are bad at reconstructing molecular graphs.
翻译:许多当代分子生成模型是分子图的变分自动编码器。其训练损失中的一项涉及重构输入,然而,最先进模型的重构能力尚未在大型且化学多样化的数据集上进行过全面比较。在这项工作中,我们表明,当几种最先进的生成模型在相同条件下评估时,它们的重构准确性出奇地低,比先前在看似更困难的数据集上报告的结果还要差。然而,我们指出,改进重构性能并不会直接带来更好的采样或优化表现。来自MoLeR模型的失败重构通常与输入相似,以不同方式组装相同的基序,并具有相似的化学性质(如溶解度)。最后,我们表明,输入分子及其失败重构通常会被不同的编码器映射到统计上可区分的后验分布,这提示后验坍缩可能无法完全解释为何VAE不擅长重构分子图。