Deep learning models for drug--target interaction (DTI) prediction often achieve strong benchmark performance without necessarily relying on mechanistically meaningful molecular features, a limitation that standard accuracy-based evaluation cannot detect. We introduce ISAAC (Intervention-based Structural Auditing Approach for Causal Reasoning), a post-hoc framework that evaluates prior-relative structural sensitivity by probing frozen models through matched mechanistic and spurious input-level interventions, independently of predictive accuracy. Applied to three sequence-based DTI architectures on the Davis benchmark, ISAAC reveals approximately 25\% relative differences in reasoning scores across models with comparable AUROC (within around 3\%), stable across training and intervention seeds and two distinct perturbation operators. These discrepancies, undetectable under conventional accuracy metrics, motivate the use of post-hoc structural auditing as a complement to standard performance evaluation in scientific machine learning for molecular modeling.
翻译:用于药物-靶点相互作用(DTI)预测的深度学习模型通常能在基准测试中取得优异性能,但未必依赖具有机制意义的分子特征——这种局限性无法通过标准化的准确性评估加以检测。我们提出ISAAC(基于干预的因果推理结构审计方法),这是一种后验框架,通过匹配的机制性干预与虚假性干预对冻结模型进行输入级探测,从而独立于预测准确性评估模型相对于先验知识的结构敏感性。在Davis基准测试中对三种基于序列的DTI架构应用ISAAC后发现,在AUROC相近(波动约3%)的模型间,推理得分存在约25%的相对差异,且该差异在不同训练种子、干预种子及两种扰动算子下保持稳定。这些在常规准确性指标下无法察觉的差异,印证了后验结构审计作为分子建模科学机器学习中标准性能评估补充手段的必要性。