Repetitive DNA (repeats) poses significant challenges for accurate and efficient genome assembly and sequence alignment. This is particularly true for metagenomic data, where genome dynamics such as horizontal gene transfer, gene duplication, and gene loss/gain complicate accurate genome assembly from metagenomic communities. Detecting repeats is a crucial first step in overcoming these challenges. To address this issue, we propose GraSSRep, a novel approach that leverages the assembly graph's structure through graph neural networks (GNNs) within a self-supervised learning framework to classify DNA sequences into repetitive and non-repetitive categories. Specifically, we frame this problem as a node classification task within a metagenomic assembly graph. In a self-supervised fashion, we rely on a high-precision (but low-recall) heuristic to generate pseudo-labels for a small proportion of the nodes. We then use those pseudo-labels to train a GNN embedding and a random forest classifier to propagate the labels to the remaining nodes. In this way, GraSSRep combines sequencing features with pre-defined and learned graph features to achieve state-of-the-art performance in repeat detection. We evaluate our method using simulated and synthetic metagenomic datasets. The results on the simulated data highlight our GraSSRep's robustness to repeat attributes, demonstrating its effectiveness in handling the complexity of repeated sequences. Additionally, our experiments with synthetic metagenomic datasets reveal that incorporating the graph structure and the GNN enhances our detection performance. Finally, in comparative analyses, GraSSRep outperforms existing repeat detection tools with respect to precision and recall.
翻译:重复DNA(重复序列)对精确高效的基因组组装和序列比对构成了重大挑战。对于宏基因组数据尤为如此,其中水平基因转移、基因复制和基因丢失/获得等基因组动态变化使得从宏基因组群落中进行准确基因组组装变得复杂。检测重复序列是克服这些挑战的关键第一步。针对这一问题,我们提出GraSSRep,一种新颖的方法,通过图神经网络(GNN)在自监督学习框架内利用组装图的结构,将DNA序列分类为重复和非重复类别。具体而言,我们将此问题定义为宏基因组组装图中的节点分类任务。以自监督方式,我们依赖高精度(但低召回率)的启发式方法为小部分节点生成伪标签,随后利用这些伪标签训练GNN嵌入和随机森林分类器,将标签传播至其余节点。通过这种方式,GraSSRep将测序特征与预定义及学习所得的图特征相结合,在重复序列检测中实现了最先进的性能。我们使用模拟和合成宏基因组数据集评估了该方法。在模拟数据上的结果突显了GraSSRep对重复序列属性的鲁棒性,证明了其在处理重复序列复杂性方面的有效性。此外,使用合成宏基因组数据集的实验表明,整合图结构和GNN增强了我们的检测性能。最后,在比较分析中,GraSSRep在精确率和召回率方面均优于现有的重复序列检测工具。