Innovations like protein diffusion have enabled significant progress in de novo protein design, which is a vital topic in life science. These methods typically depend on protein structure encoders to model residue backbone frames, where atoms do not exist. Most prior encoders rely on atom-wise features, such as angles and distances between atoms, which are not available in this context. Thus far, only several simple encoders, such as IPA, have been proposed for this scenario, exposing the frame modeling as a bottleneck. In this work, we proffer the Vector Field Network (VFN), which enables network layers to perform learnable vector computations between coordinates of frame-anchored virtual atoms, thus achieving a higher capability for modeling frames. The vector computation operates in a manner similar to a linear layer, with each input channel receiving 3D virtual atom coordinates instead of scalar values. The multiple feature vectors output by the vector computation are then used to update the residue representations and virtual atom coordinates via attention aggregation. Remarkably, VFN also excels in modeling both frames and atoms, as the real atoms can be treated as the virtual atoms for modeling, positioning VFN as a potential universal encoder. In protein diffusion (frame modeling), VFN exhibits an impressive performance advantage over IPA, excelling in terms of both designability (67.04% vs. 53.58%) and diversity (66.54% vs. 51.98%). In inverse folding (frame and atom modeling), VFN outperforms the previous SoTA model, PiFold (54.7% vs. 51.66%), on sequence recovery rate. We also propose a method of equipping VFN with the ESM model, which significantly surpasses the previous ESM-based SoTA (62.67% vs. 55.65%), LM-Design, by a substantial margin.
翻译:以蛋白质扩散为代表的创新方法推动了生命科学重要课题——从头蛋白质设计的显著进展。这类方法通常依赖蛋白质结构编码器对残基主链骨架进行建模,而该场景中原子并不存在。先前的多数编码器依赖于原子层面的特征(如原子间角度和距离),但在当前上下文中无法获取。迄今为止,只有IPA等少数简单编码器被提出用于该场景,使得骨架建模成为瓶颈。本研究提出了向量场网络(VFN),该网络层能够在骨架锚定的虚拟原子坐标之间执行可学习的向量计算,从而显著提升骨架建模能力。向量计算机制类似于线性层,每个输入通道接收3D虚拟原子坐标而非标量值。通过注意力聚合机制,向量计算输出的多个特征向量被用于更新残基表征和虚拟原子坐标。值得注意的是,VFN在骨架与原子联合建模中同样表现出色——真实原子可作为虚拟原子进行建模,使其具备成为通用编码器的潜力。在蛋白质扩散(骨架建模)任务中,VFN相较于IPA展现出显著性能优势:可设计性(67.04% vs 53.58%)和多样性(66.54% vs 51.98%)均大幅提升。在逆折叠(骨架与原子联合建模)任务中,VFN在序列恢复率指标上超越先前最优模型PiFold(54.7% vs 51.66%)。我们还提出将VFN与ESM模型结合的方法,该方法以62.67%的序列恢复率大幅超越此前基于ESM的最优模型LM-Design(55.65%)。