Biological AI models increasingly predict complex cellular responses, yet their learned representations remain disconnected from the molecular processes they aim to capture. We present CDT-III, which extends mechanism-oriented AI across the full central dogma: DNA, RNA, and protein. Its two-stage Virtual Cell Embedder architecture mirrors the spatial compartmentalization of the cell: VCE-N models transcription in the nucleus and VCE-C models translation in the cytosol. On five held-out genes, CDT-III achieves per-gene RNA r=0.843 and protein r=0.969. Adding protein prediction improves RNA performance (r=0.804 to 0.843), demonstrating that downstream tasks regularize upstream representations. Protein supervision sharpens DNA-level interpretability, increasing CTCF enrichment by 30%. Analysis of experimentally measured mRNA and protein responses reveals that the majority of genes with observable mRNA changes show opposite protein-level changes (66.7% at |log2FC|>0.01, rising to 87.5% at |log2FC|>0.02), exposing a fundamental limitation of RNA-only perturbation models. Despite this pervasive direction discordance, CDT-III correctly predicts both mRNA and protein responses. Applied to in silico CD52 knockdown approximating Alemtuzumab, the model predicts 29/29 protein changes correctly and rediscovers 5 of 7 known clinical side effects without clinical data. Gradient-based side effect profiling requires only unperturbed baseline data (r=0.939), enabling screening of all 2,361 genes without new experiments.


翻译:生物人工智能模型日益能够预测复杂的细胞反应,但其学习到的表征仍与目标捕捉的分子过程相脱节。我们提出CDT-III,将机制导向的人工智能扩展至整个中心法则:DNA、RNA与蛋白质。其双阶段虚拟细胞嵌入器架构模拟了细胞的空间区隔化:VCE-N建模细胞核内的转录过程,VCE-C建模细胞质内的翻译过程。在五个保留基因上,CDT-III实现了单基因RNA r=0.843与蛋白质r=0.969的预测性能。加入蛋白质预测后,RNA性能从r=0.804提升至0.843,表明下游任务对上游表征具有正则化作用。蛋白质监督增强了DNA层面的可解释性,使CTCF富集度提升30%。对实验测量的mRNA与蛋白质响应的分析揭示,大多数具有可观测mRNA变化的基因表现出相反的蛋白质水平变化(在|log2FC|>0.01时占66.7%,在|log2FC|>0.02时升至87.5%),暴露了仅依赖RNA扰动模型的根本局限性。尽管存在这种普遍的方向不一致性,CDT-III仍能正确预测mRNA与蛋白质响应。应用于模拟阿仑单抗作用的计算机CD52基因敲除时,模型正确预测了29/29个蛋白质变化,并在无临床数据的情况下重新发现了7种已知临床副作用中的5种。基于梯度的副作用分析只需未扰动的基线数据(r=0.939),使无需新实验即可对全部2361个基因进行筛选。

0
下载
关闭预览

相关内容

AlphaFold、人工智能(AI)和蛋白变构
专知会员服务
12+阅读 · 2022年8月28日
放弃幻想,全面拥抱Transformer:NLP三大特征抽取器(CNN/RNN/TF)比较
黑龙江大学自然语言处理实验室
10+阅读 · 2019年1月15日
类脑计算的前沿论文,看我们推荐的这7篇
人工智能前沿讲习班
21+阅读 · 2019年1月7日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
博士论文 | 面向大模型推理的内存高效算法
专知会员服务
2+阅读 · 7月27日
美空军新型反无人机部队初探
专知会员服务
7+阅读 · 7月27日
《防空交战流程的概率建模研究》
专知会员服务
10+阅读 · 7月27日
ICML 2026 教程 | 数值优化理论还重要吗?
专知会员服务
6+阅读 · 7月26日
ICM 2026 | 陶哲轩:人工智能时代的数学
专知会员服务
9+阅读 · 7月26日
《反无人机交战场景下的战斗归零研究》
专知会员服务
7+阅读 · 7月26日
博士论文 | 用代码结构感知方法推进代码大模型
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员