In the shotgun sequencing channel, the input sequence (possibly, a long DNA sequence composed of nucleotide bases) is read into multiple fragments (called `reads') of much shorter lengths. In the context of DNA data storage, the capacity of this channel was identified in a recent work, assuming that the reads themselves are noiseless substrings of the original sequence. Modern shotgun sequencers however also output quality scores for each base read, indicating the confidence in its identification. Bases with low quality scores can be considered to be erased. Motivated by this, we consider the shotgun sequencing channel with erasures, where each symbol in any read can be independently erased with some probability $\delta$. We identify achievable rates for this channel, using a random code construction and a decoder that uses typicality-like arguments to merge the reads.
翻译:在鸟枪测序信道中,输入序列(可能为由核苷酸碱基组成的长DNA序列)被读取为多个长度较短的片段(称为“读取片段”)。在DNA数据存储背景下,近期研究在假设读取片段本身是原始序列的无噪声子串的前提下,确定了该信道的容量。然而,现代鸟枪测序仪还会为每个碱基读取输出质量评分,以指示其识别的置信度。质量评分较低的碱基可视为被擦除。受此启发,我们考虑了具有擦除的鸟枪测序信道,其中任意读取片段中的每个符号可能以概率$\delta$独立地被擦除。通过采用随机码构造方法和利用类似典型性论证来合并读取片段的解码器,我们给出了该信道的可达速率。