DNA sequence classification is a fundamental task in computational biology with vast implications for applications such as disease prevention and drug design. Therefore, fast high-quality sequence classifiers are significantly important. This paper introduces ClaPIM, a scalable DNA sequence classification architecture based on the emerging concept of hybrid in-crossbar and near-crossbar memristive processing-in-memory (PIM). We enable efficient and high-quality classification by uniting the filter and search stages within a single algorithm. Specifically, we propose a custom filtering technique that drastically narrows the search space and a search approach that facilitates approximate string matching through a distance function. ClaPIM is the first PIM architecture for scalable approximate string matching that benefits from the high density of memristive crossbar arrays and the massive computational parallelism of PIM. Compared with Kraken2, a state-of-the-art software classifier, ClaPIM provides significantly higher classification quality (up to 20x improvement in F1 score) and also demonstrates a 1.8x throughput improvement. Compared with EDAM, a recently-proposed SRAM-based accelerator that is restricted to small datasets, we observe both a 30.4x improvement in normalized throughput per area and a 7% increase in classification precision.
翻译:DNA序列分类是计算生物学中的基础任务,对疾病预防和药物设计等应用具有深远影响。因此,快速高质量的序列分类器至关重要。本文提出ClaPIM,一种基于混合交叉阵列与近交叉阵列忆阻存内处理(PIM)新兴概念的可扩展DNA序列分类架构。我们通过将过滤与搜索阶段统一到单一算法中实现高效高质量分类。具体而言,我们提出一种定制化过滤技术大幅缩小搜索空间,并设计一种通过距离函数实现近似字符串匹配的搜索方法。ClaPIM是首个支持可扩展近似字符串匹配的PIM架构,其优势源于忆阻交叉阵列的高密度特性与PIM的大规模计算并行性。与当前最先进的软件分类器Kraken2相比,ClaPIM在分类质量上显著提升(F1分数最高提升20倍),同时吞吐量提高1.8倍。与近期提出的受限于小数据集的基于SRAM的加速器EDAM相比,我们观察到归一化每面积吞吐量提升30.4倍,分类精确度提高7%。