The Hayashi-Yoshida (\HY)-estimator exhibits an intrinsic, telescoping property that leads to an often overlooked computational bias, which we denote,formulaic or intrinsic bias. This formulaic bias results in data loss by cancelling out potentially relevant data points, the nonextant data points. This paper attempts to formalize and quantify the data loss arising from this bias. In particular, we highlight the existence of nonextant data points via a concrete example, and prove necessary and sufficient conditions for the telescoping property to induce this type of formulaic bias.Since this type of bias is nonexistent when inputs, i.e., observation times, $\Pi^{(1)} :=(t_i^{(1)})_{i=0,1,\ldots}$ and $\Pi^{(2)} :=(t_j^{(2)})_{j=0,1,\ldots}$, are synchronous, we introduce the (a,b)-asynchronous adversary. This adversary generates inputs $\Pi^{(1)}$ and $\Pi^{(2)}$ according to two independent homogenous Poisson processes with rates a>0 and b>0, respectively. We address the foundational questions regarding cumulative minimal (or least) average data point loss, and determine the values for a and b. We prove that for equal rates a=b, the minimal average cumulative data loss over both inputs is attained and amounts to 25\%. We present an algorithm, which is based on our theorem, for computing the exact number of nonextant data points given inputs $\Pi^{(1)}$ and $\Pi^{(2)}$, and suggest alternative methods. Finally, we use simulated data to empirically compare the (cumulative) average data loss of the (\HY)-estimator.
翻译:Hayashi-Yoshida (\HY)估计量具有一种内在的伸缩性质,会导致一种常被忽视的计算偏差,我们将其称为公式化偏差或内在偏差。这种公式化偏差通过消除潜在相关数据点(即不存在的数据点)导致数据丢失。本文试图正式定义并量化由该偏差引起的数据丢失。具体而言,我们通过具体实例强调不存在数据点的存在,并证明伸缩性质引发此类公式化偏差的充要条件。由于当输入(即观测时间点)$\Pi^{(1)} :=(t_i^{(1)})_{i=0,1,\ldots}$ 和 $\Pi^{(2)} :=(t_j^{(2)})_{j=0,1,\ldots}$ 同步时此类偏差不存在,我们引入(a,b)-异步对抗模型。该对抗模型根据两个独立同质泊松过程分别生成输入$\Pi^{(1)}$和$\Pi^{(2)}$,其速率分别为a>0和b>0。我们探讨关于累积最小(或最少)平均数据点丢失的基础性问题,并确定a与b的取值。我们证明当速率相等a=b时,两个输入上的最小平均累积数据丢失可达25%。基于该定理,我们提出一种算法,用于在给定输入$\Pi^{(1)}$和$\Pi^{(2)}$时计算不存在数据点的精确数量,并建议替代方法。最后,我们使用模拟数据对(\HY)估计量的(累积)平均数据丢失进行实证比较。