Binding prediction models accelerate therapeutic antibody and TCR discovery, but their performance on new datasets is unpredictable, often leading to low discovery rates. Density-ratio methods (PAPE, M-CBPE) provide label-free performance estimation for binary classification, but their assumptions and aggregate-only outputs limit binding prediction on neoepitopes, antigen variants and chemical scaffolds. Here we present CaliPPer (Calibration and Prediction of Performance), a post-hoc framework pairing a multi-chain Sample-to-Domain Distance (S2DD) with distance-aware Bayesian recalibration, operating at three resolutions: generalisability score, aggregate performance prediction, and per-sample confidence. Across ten models, eight architectures and two immune-receptor domains, CaliPPer attains distance--performance correlations $|r|=0.80\text{--}0.92$, predicts AUROC/AP/F1 with mean absolute errors $0.008\text{--}0.070$, and improves AUROC by up to $+0.20$ on unseen epitopes/variants. Applied retrospectively to five published TCR, BCR, MHC--peptide and small-molecule studies, CaliPPer raises true discovery rates in all five (e.g.\ $0/5 \to 3/5$ confirmed neoantigens), providing a triage layer between computational prediction and experimental validation.
翻译:结合预测模型可加速治疗性抗体和TCR的发现,但其在新数据集上的性能难以预测,常导致发现率低下。密度比方法(PAPE、M-CBPE)为二分类任务提供了无标签的性能估计,但其假设条件和仅输出聚合结果的特性限制了其在新生抗原表位、抗原变体和化学骨架结合预测中的应用。本文提出CaliPPer(性能校准与预测),这是一种事后框架,将多链样本到域距离(S2DD)与距离感知的贝叶斯校准相结合,在三个分辨率层面运行:泛化能力评分、聚合性能预测和逐样本置信度。在十个模型、八种架构和两个免疫受体域上,CaliPPer实现了距离与性能之间的相关性$|r|=0.80\text{--}0.92$,AUROC/AP/F1的预测平均绝对误差为$0.008\text{--}0.070$,并在未见过的抗原表位/变体上将AUROC提升高达$+0.20$。将CaliPPer回顾性应用于五篇已发表的TCR、BCR、MHC-肽及小分子研究中,它成功提高了所有五项研究中的真实发现率(例如,$0/5 \to 3/5$ 确认的新抗原),在计算预测与实验验证之间提供了筛选层。