The concept of dimension is essential to grasp the complexity of data. A naive approach to determine the dimension of a dataset is based on the number of attributes. More sophisticated methods derive a notion of intrinsic dimension (ID) that employs more complex feature functions, e.g., distances between data points. Yet, many of these approaches are based on empirical observations, cannot cope with the geometric character of contemporary datasets, and do lack an axiomatic foundation. A different approach was proposed by V. Pestov, who links the intrinsic dimension axiomatically to the mathematical concentration of measure phenomenon. First methods to compute this and related notions for ID were computationally intractable for large-scale real-world datasets. In the present work, we derive a computationally feasible method for determining said axiomatic ID functions. Moreover, we demonstrate how the geometric properties of complex data are accounted for in our modeling. In particular, we propose a principle way to incorporate neighborhood information, as in graph data, into the ID. This allows for new insights into common graph learning procedures, which we illustrate by experiments on the Open Graph Benchmark.
翻译:维度概念对于理解数据的复杂性至关重要。确定数据集维度的朴素方法基于属性数量。更复杂的方法通过利用更复杂的特征函数(例如数据点之间的距离)来推导本征维度(ID)的概念。然而,许多此类方法基于经验观察,无法应对当代数据集的几何特性,且缺乏公理化基础。V. Pestov提出了一种不同的方法,该方法将本征维度与数学中的测度集中现象进行公理化关联。计算这种ID及相关概念的首批方法在大规模真实世界数据集上缺乏计算可行性。在本工作中,我们推导出一种计算可行的方法来确定所述公理化ID函数。此外,我们展示了复杂数据的几何特性如何在我们的建模中得到体现。特别地,我们提出了一种将邻域信息(如图数据中的邻域)融入ID的原则性方法。这为理解常见的图学习过程提供了新的视角,我们通过Open Graph Benchmark上的实验加以说明。