The problem of distributed matrix-vector product is considered, where the server distributes the task of the computation among $n$ worker nodes, out of which $L$ are compromised (but non-colluding) and may return incorrect results. Specifically, it is assumed that the compromised workers are unreliable, that is, at any given time, each compromised worker may return an incorrect and correct result with probabilities $\alpha$ and $1-\alpha$, respectively. Thus, the tests are noisy. This work proposes a new probabilistic group testing approach to identify the unreliable/compromised workers with $O\left(\frac{L\log(n)}{\alpha}\right)$ tests. Moreover, using the proposed group testing method, sparse parity-check codes are constructed and used in the considered distributed computing framework for encoding, decoding and identifying the unreliable workers. This methodology has two distinct features: (i) the cost of identifying the set of $L$ unreliable workers at the server can be shown to be considerably lower than existing distributed computing methods, and (ii) the encoding and decoding functions are easily implementable and computationally efficient.
翻译:考虑分布式矩阵-向量乘积问题,其中服务器将计算任务分配给$n$个工作节点,其中$L$个节点被攻陷(但非合谋)且可能返回错误结果。具体而言,假设被攻陷的工作节点不可靠,即在任意时刻,每个被攻陷节点以概率$\alpha$返回错误结果,以概率$1-\alpha$返回正确结果,因此测试过程带有噪声。本研究提出一种新的概率组测试方法,通过$O\left(\frac{L\log(n)}{\alpha}\right)$次测试即可识别不可靠/被攻陷的工作节点。此外,利用所提出的组测试方法,构建了稀疏奇偶校验码,并将其应用于所考虑的分布式计算框架中,实现编码、解码及不可靠工作节点识别。该方法具有两个显著特征:(i)与现有分布式计算方法相比,服务器识别$L$个不可靠工作节点集合的成本显著降低;(ii)编码与解码函数易于实现且计算效率高。