Measuring geometric similarity between high-dimensional network representations is a topic of longstanding interest to neuroscience and deep learning. Although many methods have been proposed, only a few works have rigorously analyzed their statistical efficiency or quantified estimator uncertainty in data-limited regimes. Here, we derive upper and lower bounds on the worst-case convergence of standard estimators of shape distance$\unicode{x2014}$a measure of representational dissimilarity proposed by Williams et al. (2021).These bounds reveal the challenging nature of the problem in high-dimensional feature spaces. To overcome these challenges, we introduce a new method-of-moments estimator with a tunable bias-variance tradeoff. We show that this estimator achieves substantially lower bias than standard estimators in simulation and on neural data, particularly in high-dimensional settings. Thus, we lay the foundation for a rigorous statistical theory for high-dimensional shape analysis, and we contribute a new estimation method that is well-suited to practical scientific settings.
翻译:测量高维网络表征之间的几何相似性是神经科学和深度学习领域长期关注的问题。尽管已有多种方法被提出,但仅有少数研究严格分析了其统计效率或量化了数据有限情境下的估计器不确定性。本文针对Williams等人(2021)提出的表征相异性度量——形状距离的标准估计器,推导了其最坏情况收敛性的上下界。这些界限揭示了高维特征空间中该问题的艰巨性。为克服这些挑战,我们引入了一种具有可调偏差-方差权衡的新型矩估计方法。实验表明,该估计器在模拟实验和神经数据中(尤其在高维场景下)相比标准估计器实现了显著更低的偏差。因此,我们为高维形状分析的严格统计理论奠定了基础,并贡献了一种适用于实际科研场景的新估计方法。