Nanomaterial-protein interactions (NPI) are pivotal to realizing the therapeutic and diagnostic potential of nanomaterials. Although AI promises to accelerate mechanistic understanding and enable rational nanomaterial design, robust generalization to unseen nanomaterials or proteins remains unresolved. Here, we present CuMMI (curriculum-guided multimodal interaction model), a generalizable, explainable, and transferable model designed to infer NPI across complex biological settings. CuMMI leverages a self-constructed million-scale NPI dataset and adopts a multi-stage curriculum centered on human plasma, with progressively broader biofluid exposure to enhance data coverage and generalizability. By integrating protein sequence, structure, and a text-encoded experimental context of 37 features, CuMMI captures complementary material-specific, biochemical, and environmental information. Sample-level quality weights are assigned to ensure full utilization of available data while mitigating low-confidence and sparsely recorded entries. Ablation studies highlight the most influential tabular features, clarifying their contribution to the prediction. Through rigorous external validation across independence-preserving temporal, nanomaterial-held-out, and protein-held-out evaluations, our framework consistently achieves good performance (mean of five classification metrics exceeding 0.75), highlighting its robustness and generalizability to unseen data. Furthermore, fine-tuning on independent gold-nanoparticle data and a held-out protein subset further delivers better performance than training from scratch with substantially fewer samples. Together, our approach enables generalizable and transferable NPI prediction and may accelerate in vitro research and applications of nanomaterials.
翻译:摘要:纳米材料-蛋白质相互作用(NPI)是实现纳米材料治疗与诊断潜力的关键。尽管人工智能有望加速机制理解并实现理性纳米材料设计,但对未见纳米材料或蛋白质的鲁棒泛化问题仍未解决。本文提出了一种具有泛化性、可解释性和迁移性的CuMMI(课程引导多模态相互作用模型),旨在推演复杂生物环境下的NPI。CuMMI利用自构建的百万级NPI数据集,采用以人血浆为核心的多阶段课程,逐步扩大生物流体暴露范围以提升数据覆盖度和泛化能力。通过整合蛋白质序列、结构及包含37个特征的文本编码实验环境信息,CuMMI捕获了互补的材料特异性、生物化学与环境信息。模型分配样本级质量权重以充分利用现有数据,同时抑制低置信度与稀疏记录条目。消融研究揭示了最具影响力的表格特征,阐释了其对预测的贡献。通过时间独立性、纳米材料留出和蛋白质留出评估的严格外部验证,本框架持续取得优异性能(五项分类指标均值超过0.75),凸显其对未见数据的鲁棒性与泛化能力。此外,在独立金纳米颗粒数据和留出蛋白质子集上的微调,以显著更少的样本取得了优于从头训练的性能。综上,本方法实现了可泛化与可迁移的NPI预测,有望加速纳米材料的体外研究与应用。