A covariance matrix estimator using two bits per entry was recently developed by Dirksen, Maly and Rauhut [Annals of Statistics, 50(6), pp. 3538-3562]. The estimator achieves near minimax rate for general sub-Gaussian distributions, but also suffers from two downsides: theoretically, there is an essential gap on operator norm error between their estimator and sample covariance when the diagonal of the covariance matrix is dominated by only a few entries; practically, its performance heavily relies on the dithering scale, which needs to be tuned according to some unknown parameters. In this work, we propose a new 2-bit covariance matrix estimator that simultaneously addresses both issues. Unlike the sign quantizer associated with uniform dither in Dirksen et al., we adopt a triangular dither prior to a 2-bit quantizer inspired by the multi-bit uniform quantizer. By employing dithering scales varying across entries, our estimator enjoys an improved operator norm error rate that depends on the effective rank of the underlying covariance matrix rather than the ambient dimension, thus closing the theoretical gap. Moreover, our proposed method eliminates the need of any tuning parameter, as the dithering scales are entirely determined by the data. Experimental results under Gaussian samples are provided to showcase the impressive numerical performance of our estimator. Remarkably, by halving the dithering scales, our estimator oftentimes achieves operator norm errors less than twice of the errors of sample covariance.
翻译:Dirksen、Maly和Rauhut [《统计学年鉴》,第50卷第6期,第3538-3562页] 近期提出了一种每项仅用两比特的协方差矩阵估计器。该估计器在一般亚高斯分布下达到了近乎极小化最优的收敛速率,但同时存在两个缺陷:理论上,当协方差矩阵的对角线仅由少数元素主导时,其估计器与样本协方差之间的算子范数误差存在本质差距;实践中,其性能严重依赖于抖动尺度,而该尺度需根据某些未知参数进行调优。本文提出了一种新的两比特协方差矩阵估计器,同时解决了上述两个问题。与Dirksen等人采用的均匀抖动符号量化器不同,我们受多比特均匀量化器的启发,在双比特量化器前采用三角抖动。通过采用随元素变化的抖动尺度,我们的估计器获得了改进的算子范数误差率,该误差率取决于底层协方差矩阵的有效秩而非环境维度,从而弥补了理论差距。此外,由于抖动尺度完全由数据决定,我们提出的方法无需任何调优参数。我们提供了高斯样本下的实验结果,展示了所提估计器优异的数值性能。值得注意的是,通过将抖动尺度减半,我们的估计器通常能实现低于样本协方差误差两倍的算子范数误差。