Progress in genomic foundation models is difficult to assess due to fragmented benchmarks, incompatible evaluation protocols, and task-specific reporting. As a result, claims of superiority or generality across models are often not directly comparable. We introduce GENEB, a large-scale diagnostic benchmark that evaluates frozen representations from 40 genomic foundation models across 100 tasks spanning 13 functional categories under a unified probing-based protocol, including few-shot regimes. GENEB enables controlled comparison across model scale, architecture, tokenization, and pretraining data while explicitly exposing task-level trade-offs. Our analysis shows that aggregate leaderboards are unstable: model rankings vary sharply across task categories, scale provides only modest and inconsistent gains, and architectural and pretraining alignment frequently outweigh parameter count. These results highlight limitations of current evaluation practices and position GENEB as a reference framework for principled comparison and category-aware model selection in genomic machine learning.
翻译:在基因组基础模型的研究进展中,由于基准测试分散、评估协议不兼容以及任务特异性报告等原因,对模型进展的评估变得困难重重。结果,关于模型优越性或泛化能力的宣称往往无法直接进行比较。我们提出了GENEB——一个大规模诊断性基准测试,它在统一的基于探针的评估协议下(包含小样本场景),对40个基因组基础模型在跨越13个功能类别的100项任务中的冻结表示进行评估。GENEB能够在明确暴露任务层面权衡的同时,对模型规模、架构、分词方法和预训练数据实现受控比较。我们的分析表明,聚合排行榜并不稳定:模型排名在不同任务类别间剧烈波动;模型规模带来的提升有限且不一致;而架构和预训练对齐性往往比参数数量更为重要。这些结果突显了当前评估实践的局限性,并将GENEB定位为基因组机器学习中原则性比较和类别感知模型选择的参考框架。