We consider error-correcting coding for DNA-based storage. We model the DNA storage channel as a multi-draw IDS channel where the input data is chunked into $M$ short DNA strands, which are copied a random number of times, and the channel outputs a random selection of $N$ noisy DNA strands. The retrieved DNA strands are prone to insertion, deletion, and substitution (IDS) errors. We propose an index-based concatenated coding scheme consisting of the concatenation of an outer code, an index code, and an inner synchronization code, where the latter two tackle IDS errors. We further propose a mismatched joint index-synchronization code maximum a posteriori probability decoder with optional clustering to infer symbolwise a posterior probabilities for the outer decoder. We compute achievable information rates for the outer code and present Monte-Carlo simulations for information-outage probabilities and frame error rates on synthetic and experimental data, respectively.
翻译:我们研究了基于DNA存储的纠错编码问题。将DNA存储信道建模为多次读取的插入-删除-替换(IDS)信道,其中输入数据被分割成$M$个短DNA链,这些链被随机复制若干次,信道输出$N$个随机选取的带噪DNA链。检索到的DNA链易发生插入、删除和替换(IDS)错误。我们提出了一种基于索引的级联编码方案,该方案由外码、索引码和内同步码的级联构成,其中后两者用于处理IDS错误。进一步地,我们提出了一种带可选聚类的失配联合索引-同步码最大后验概率译码器,用于推断外码译码的逐符号后验概率。我们计算了外码的可达信息率,并通过蒙特卡洛仿真分别分析了合成数据和实验数据上的信息中断概率和帧错误率。