Scatter plots are widely recognized as fundamental tools for illustrating the relationship between two numerical variables. Despite this, based on solid theoretical foundations, scatter plots generated from pairs of continuous random variables may not serve as reliable tools for assessing dependence. Sklar's Theorem implies that scatter plots created from ranked data are preferable for such analysis as they exclusively convey information pertinent to dependence. This is in stark contrast to conventional scatter plots, which also encapsulate information about the variables' marginal distributions. Such additional information is extraneous to dependence analysis and can obscure the visual interpretation of the variables' relationship. In this article, we delve into the theoretical underpinnings of these ranked data scatter plots, hereafter referred to as rank plots. We offer insights into interpreting the information they reveal and examine their connections with various association measures, including Pearson's and Spearman's correlation coefficients, as well as Schweizer-Wolff's measure of dependence. Furthermore, we introduce a novel graphical combination for dependence analysis, termed a dplot, and demonstrate its efficacy through real data examples.
翻译:散点图被广泛认为是展示两个数值变量之间关系的基本工具。然而,基于坚实的理论基础,由连续型随机变量对生成的散点图并非评估相关性的可靠工具。Sklar定理表明,基于排序数据生成的散点图更适用于此类分析,因为它们仅传递与相关性相关的信息。这与传统散点图形成鲜明对比,后者还包含变量边缘分布的信息。这些额外信息与相关性分析无关,且会模糊变量关系的视觉解读。本文深入探讨了此类排序数据散点图(以下简称秩图)的理论基础,并提供了解读其所揭示信息的见解,同时考察了其与多种关联度量(包括皮尔逊相关系数、斯皮尔曼相关系数以及Schweizer-Wolff相关性度量)之间的联系。此外,我们提出了一种新的相关性分析图形组合——d图,并通过实际数据案例验证了其有效性。