Similarity index is an important scientific tool frequently used to determine whether different pairs of entities are similar with respect to some prefixed characteristics. Some standard measures of similarity index include Jaccard index, S{\o}rensen-Dice index, and Simpson's index. Recently, a better index ($\hat{\alpha}$) for the co-occurrence and/or similarity has been developed, and this measure really outperforms and gives theoretically supported reasonable predictions. However, the measure $\hat{\alpha}$ is not data dependent. In this article we propose a new measure of similarity which depends strongly on the data before introducing randomness in prevalence. Then, we propose a new method of randomization which changes the whole pattern of results. Before randomization our measure is similar to the Jaccard index, while after randomization it is close to $\hat{\alpha}$. We consider the popular ecological dataset from the Tuscan Archipelago, Italy; and compare the performance of the proposed index to other measures. Since our proposed index is data dependent, it has some interesting properties which we illustrate in this article through numerical studies.
翻译:相似性指数是一种重要的科学工具,常用于判断不同实体对在若干预设特征上是否相似。标准的相似性指数测度包括Jaccard指数、Sørensen-Dice指数和Simpson指数。近期,针对共现和/或相似性,已开发出更优的指数($\hat{\alpha}$),该测度在实际应用中表现卓越,并能提供具有理论支撑的合理预测。然而,$\hat{\alpha}$测度并不依赖于数据。本文提出一种新的相似性测度,该测度在引入流行率随机性之前高度依赖数据。随后,我们提出一种新的随机化方法,该方法能够彻底改变结果模式。在随机化之前,我们的测度与Jaccard指数相似;而随机化后,其值接近$\hat{\alpha}$。我们采用来自意大利托斯卡纳群岛的经典生态数据集,并将所提指数与其他测度的性能进行比较。由于所提指数具有数据依赖性,它展现出若干有趣特性,本文将通过数值研究对其进行阐述。