We investigate the fundamental limits of the recently proposed random access coverage depth problem for DNA data storage. Under this paradigm, it is assumed that the user information consists of $k$ information strands, which are encoded into $n$ strands via some generator matrix $G$. In the sequencing process, the strands are read uniformly at random, since each strand is available in a large number of copies. In this context, the random access coverage depth problem refers to the expected number of reads (i.e., sequenced strands) until it is possible to decode a specific information strand, which is requested by the user. The goal is to minimize the maximum expectation over all possible requested information strands, and this value is denoted by $T_{\max}(G)$. This paper introduces new techniques to investigate the random access coverage depth problem, which capture its combinatorial nature. We establish two general formulas to find $T_{max}(G)$ for arbitrary matrices. We introduce the concept of recovery balanced codes and combine all these results and notions to compute $T_{\max}(G)$ for MDS, simplex, and Hamming codes. We also study the performance of modified systematic MDS matrices and our results show that the best results for $T_{\max}(G)$ are achieved with a specific mix of encoded strands and replication of the information strands.
翻译:我们研究了最近提出的DNA数据存储中随机访问覆盖深度问题的基本极限。在该范式下,假设用户信息由$k$条信息链组成,这些链通过某个生成矩阵$G$编码为$n$条链。在测序过程中,由于每条链有大量副本,链被均匀随机地读取。在此背景下,随机访问覆盖深度问题指在能够解码用户请求的特定信息链之前所需的预期读取次数(即测序链数)。目标是使所有可能请求的信息链上的最大期望值最小化,该值记为$T_{\max}(G)$。本文引入新方法来研究随机访问覆盖深度问题,捕捉其组合本质。我们建立了两个通用公式,用于计算任意矩阵的$T_{\max}(G)$。我们引入了恢复平衡码的概念,并结合所有结果和概念,计算了MDS码、单纯码和汉明码的$T_{\max}(G)$。我们还研究了改进的系统性MDS矩阵的性能,结果表明,通过编码链与信息链复制的特定混合方式,可获得$T_{\max}(G)$的最优结果。