Foundation models in genomics have shown mixed success compared to their counterparts in natural language processing. Yet, the reasons for their limited effectiveness remain poorly understood. In this work, we investigate the role of entropy as a fundamental factor limiting the capacities of such models to learn from their training data and develop foundational capabilities. We train ensembles of models on text and DNA sequences and analyze their predictions, static embeddings, and empirical Fisher information flow. We show that the high entropy of genomic sequences -- from the point of view of unseen token prediction -- leads to near-uniform output distributions, disagreement across models, and unstable static embeddings, even for models that are matched in architecture, training and data. We then demonstrate that models trained on DNA concentrate Fisher information in embedding layers, seemingly failing to exploit inter-token relationships. Our results suggest that self-supervised training from sequences alone may not be applicable to genomic data, calling into question the assumptions underlying current methodologies for training genomic foundation models.
翻译:基因组学中的基础模型相较于自然语言处理领域的同类模型表现参差不齐。然而,其有效性受限的根本原因尚未明晰。本研究探讨了熵作为限制此类模型从训练数据中学习并发展基础能力的关键因素所扮演的角色。我们通过在文本和DNA序列上训练模型集成,分析其预测结果、静态嵌入以及经验费希尔信息流。研究表明,从未知词元预测的角度看,基因组序列的高熵特性会导致近似均匀的输出分布、模型间的预测分歧以及不稳定的静态嵌入——即便模型在架构、训练方法和数据规模方面完全匹配。进一步研究发现,基于DNA训练的模型将费希尔信息集中于嵌入层,似乎未能有效利用词元间的关联关系。我们的结果表明,仅依赖序列本身的自监督训练方法可能不适用于基因组数据,这质疑了当前用于训练基因组基础模型的方法论假设。