The genome sequence contains the blueprint for governing cellular processes. While the availability of genomes has vastly increased over the last decades, experimental annotation of the various functional, non-coding and regulatory elements encoded in the DNA sequence remains both expensive and challenging. This has sparked interest in unsupervised language modeling of genomic DNA, a paradigm that has seen great success for protein sequence data. Although various DNA language models have been proposed, evaluation tasks often differ between individual works, and might not fully recapitulate the fundamental challenges of genome annotation, including the length, scale and sparsity of the data. In this study, we introduce BEND, a Benchmark for DNA language models, featuring a collection of realistic and biologically meaningful downstream tasks defined on the human genome. We find that embeddings from current DNA LMs can approach performance of expert methods on some tasks, but only capture limited information about long-range features. BEND is available at https://github.com/frederikkemarin/BEND.
翻译:摘要:基因组序列包含调控细胞过程的蓝图。尽管近几十年来基因组的可用性大幅增加,但对DNA序列中编码的各种功能元件、非编码元件和调控元件的实验注释仍然既昂贵又具有挑战性。这引发了人们对基因组DNA无监督语言建模的兴趣,这一范式在蛋白质序列数据上已取得巨大成功。尽管已提出多种DNA语言模型,但不同研究的评估任务往往存在差异,且可能无法完全体现基因组注释的基本挑战,包括数据的长度、尺度和稀疏性。在本研究中,我们提出BEND(DNA语言模型基准测试),包含一组在人类基因组上定义的真实且具有生物学意义的下游任务。我们发现,当前DNA语言模型生成的嵌入在某些任务上可接近专家方法的性能,但仅能捕获有限的长程特征信息。BEND代码已开源,访问地址为https://github.com/frederikkemarin/BEND。