Mechanism-level drug-drug interaction (DDI) prediction requires identifying which enzyme or pharmacodynamic axis is implicated, in which direction, and with which evidence -- not merely whether two drugs interact. We introduce a reproducible mechanism-level DDI labelling and evaluation protocol with a structured 7-family/147-subtype taxonomy, leakage-safe cold-split protocols, and auditable reasoning metrics for evaluating pharmacological prediction beyond flat interaction classification. We propose a pipeline that produces a 7B reasoning MARD (Mirror-Augmented Reasoning Distillation), combining three training innovations: a single-token KL divergence on direction tag that ties the model's prediction, per-loss PRM-weighted DPO with programmatic hard negatives, and a leakage-safe mechanism-aware retrieval channel. Process-reward step labels are automatically verifiable against DrugBank-structured fields, requiring no human or LLM judges. On the April-2026 DrugBank release, our MARD-7B is the only system in a 32-system comparison whose accuracy survives drug-pair novelty, beating the best baseline by +13.9 pp and GPT-4o by +6.7 pp at ~1% of frontier API cost. Further analysis reveals an anti-memorisation signature where accuracy improves on rarely seen drugs, suggesting that gain comes from structured pharmacological reasoning rather than drug-frequency memorisation. We release corpus, DDI-PRM, retrieval index, and training code.
翻译:摘要:机制级药物相互作用(DDI)预测需要识别涉及何种酶或药效轴、作用方向及相应证据——而不仅仅是判断两种药物是否发生相互作用。我们提出了一种可复现的机制级DDI标注与评估协议,包含结构化的7家族/147亚型分类体系、防泄漏的冷分割协议,以及用于评估药理学预测(超越平面相互作用分类)的可审计推理指标。我们构建了一个可生成7B级推理模型MARD(镜像增强推理蒸馏)的流水线,该模型融合了三项训练创新:在方向标签上使用单token KL散度以约束模型预测方向、基于逐损失PRM加权DPO结合程序化硬负样本,以及防泄漏的机制感知检索通道。过程奖励步骤标签可直接依据DrugBank结构化字段进行自动验证,无需人工或LLM评判员。在2026年4月版DrugBank数据上,我们的MARD-7B是32个系统对比中唯一一个在药物对新颖性条件下仍保持准确率的系统,以约1%的先进API成本,比最优基线模型高出13.9个百分点,比GPT-4o高出6.7个百分点。进一步分析揭示了反记忆特征:准确率在罕见药物上反而提升,表明性能提升源于结构化药理学推理而非药物频率记忆。我们公开了语料库、DDI-PRM、检索索引及训练代码。