In many scientific research, it is often imperative to determine whether pairs of entities have similarities in themselves or not. There are standard approaches to this problem, such as Jaccard, Sorensen Dice, and Simpson. Recently, a better index for the analysis of cooccurrence and similarity was developed and it reversed all the results obtained by standard indices and supported theoretical predictions. In this paper, we propose a new method of similarity using MLE, PCA, LDA, and clustering. Our index depends strongly on the data before introducing randomness in prevalence. Then we propose a new method of randomization which changed the whole pattern of the results. Before randomization, it was strongly dependent o the prevalence and hence was following the pattern of the Jaccard index. So, we introduce the new randomization technique, and hence the whole results reversed and followed that of alpha. Also, we will show some limitations of alpha which we try to resolve through different pathways.
翻译:在众多科学研究中,常常需要判断实体对之间是否存在内在相似性。针对这一问题已有多种标准方法,如杰卡德指数、索伦森-戴斯指数和辛普森指数。近期,一种新的共现与相似性分析指数被开发出来,该指数颠覆了所有标准方法所得结果,并支持了理论预测。本文提出一种基于极大似然估计、主成分分析、线性判别分析和聚类的新型相似度计算方法。该指数在引入随机性之前高度依赖于数据本身,因此我们进一步提出了一种全新的随机化方法,该方法完全改变了结果模式。随机化前,该指数与出现频率高度相关,遵循杰卡德指数的变化规律;而通过引入新的随机化技术,整体结果发生逆转,转而遵循α指数的变化规律。同时,我们将揭示α指数存在的若干局限性,并尝试通过不同途径加以解决。