The amount of sequencing data for SARS-CoV-2 is several orders of magnitude larger than any virus. This will continue to grow geometrically for SARS-CoV-2, and other viruses, as many countries heavily finance genomic surveillance efforts. Hence, we need methods for processing large amounts of sequence data to allow for effective yet timely decision-making. Such data will come from heterogeneous sources: aligned, unaligned, or even unassembled raw nucleotide or amino acid sequencing reads pertaining to the whole genome or regions (e.g., spike) of interest. In this work, we propose \emph{ViralVectors}, a compact feature vector generation from virome sequencing data that allows effective downstream analysis. Such generation is based on \emph{minimizers}, a type of lightweight "signature" of a sequence, used traditionally in assembly and read mapping -- to our knowledge, the first use minimizers in this way. We validate our approach on different types of sequencing data: (a) 2.5M SARS-CoV-2 spike sequences (to show scalability); (b) 3K Coronaviridae spike sequences (to show robustness to more genomic variability); and (c) 4K raw WGS reads sets taken from nasal-swab PCR tests (to show the ability to process unassembled reads). Our results show that ViralVectors outperforms current benchmarks in most classification and clustering tasks.
翻译:SARS-CoV-2的测序数据量比其他任何病毒高出数个数量级。随着许多国家对基因组监测工作的大规模投入,这一数据量对于SARS-CoV-2及其他病毒而言将持续呈几何级增长。因此,我们需要能够处理海量序列数据的方法,以支持及时有效的决策制定。这些数据将来自异质来源:涵盖全基因组或特定区域(如刺突蛋白)的已比对、未比对甚至未组装的原始核苷酸或氨基酸测序读段。本文提出ViralVectors——一种从病毒组测序数据中生成紧凑特征向量的方法,可支持高效的下游分析。该生成方法基于minimizers(一种轻量级序列"签名"),该工具传统上用于序列组装和读段比对——据我们所知,这是首次将minimizers以这种方式应用。我们在三类不同的测序数据上验证了该方法:(a) 250万条SARS-CoV-2刺突蛋白序列(以展示可扩展性);(b) 3000条冠状病毒科刺突蛋白序列(以展示对更大基因组变异的鲁棒性);以及(c) 4000组来自鼻拭子PCR检测的原始全基因组测序读段(以展示处理未组装读段的能力)。结果表明,在大多数分类和聚类任务中,ViralVectors的表现优于现有基准方法。