Due to the high mutation rate of the virus, the COVID-19 pandemic evolved rapidly. Certain variants of the virus, such as Delta and Omicron, emerged with altered viral properties leading to severe transmission and death rates. These variants burdened the medical systems worldwide with a major impact to travel, productivity, and the world economy. Unsupervised machine learning methods have the ability to compress, characterize, and visualize unlabelled data. This paper presents a framework that utilizes unsupervised machine learning methods to discriminate and visualize the associations between major COVID-19 variants based on their genome sequences. These methods comprise a combination of selected dimensionality reduction and clustering techniques. The framework processes the RNA sequences by performing a k-mer analysis on the data and further visualises and compares the results using selected dimensionality reduction methods that include principal component analysis (PCA), t-distributed stochastic neighbour embedding (t-SNE), and uniform manifold approximation projection (UMAP). Our framework also employs agglomerative hierarchical clustering to visualize the mutational differences among major variants of concern and country-wise mutational differences for selected variants (Delta and Omicron) using dendrograms. We also provide country-wise mutational differences for selected variants via dendrograms. We find that the proposed framework can effectively distinguish between the major variants and has the potential to identify emerging variants in the future.
翻译:由于病毒的高突变率,COVID-19疫情迅速演变。德尔塔和奥密克戎等特定变异株的出现改变了病毒特性,导致严重的传播率和死亡率。这些变异株给全球医疗系统带来沉重负担,并对旅行、生产力和世界经济造成重大影响。无监督机器学习方法能够压缩、表征和可视化未标记数据。本文提出一个利用无监督机器学习方法基于基因组序列区分和可视化主要COVID-19变异株之间关联性的框架。这些方法包含选定的降维与聚类技术的组合。该框架通过对RNA序列进行k-mer分析,并利用选定的降维方法(包括主成分分析(PCA)、t分布随机邻域嵌入(t-SNE)和统一流形逼近与投影(UMAP))进一步可视化和比较结果。我们的框架还采用凝聚层次聚类,通过树状图可视化主要关切变异株之间的突变差异以及选定变异株(德尔塔和奥密克戎)按国家划分的突变差异。我们通过树状图提供了选定变异株按国家划分的突变差异。研究发现,所提出的框架能有效区分主要变异株,并具有未来识别新兴变异株的潜力。