Synthetic DNA approaches 227.5 exabytes per gram of storage density with stability over millennial timescales. Realising this capacity requires error-correction codes that recover data from substantial synthesis and sequencing errors. Existing codecs convert noisy sequencer output into discrete base calls before error correction, discarding probabilistic information about which positions are reliable. Here we present a coding scheme that retains the sequencer's per-position posterior distributions through an integrated decoder of profile hidden Markov model alignment, log-product fusion across reads, and ordered-statistics decoding. On the DT4DDS channel simulator, the codec recovers 155.8 and 25.9 exabytes per gram of dsDNA under high- and low-fidelity conditions, exceeding the highest prior-art density on each channel by 11 and 52 percent. Under a single-encode-then-degrade protocol mapped to depurination kinetics at 25 °C in the dry state, the codec projects 282 years of decodable storage at 17.1 exabytes per gram. These results place DNA storage density within reach of the Shannon bound of the underlying channel.
翻译:合成DNA的存储密度可达每克227.5 EB(艾字节),且能保持千年时间尺度的稳定性。实现这一容量需要能够从大量合成和测序错误中恢复数据的纠错码。现有编解码器在纠错前将含噪声的测序输出转换为离散碱基调用,丢弃了关于哪些位置可靠的概率信息。本文提出一种编码方案,通过集成剖面隐马尔可夫模型比对、跨读段log乘积融合和有序统计解码,保留测序仪每个位置的后验分布。在DT4DDS通道模拟器上,该编解码器在高保真和低保真条件下分别恢复每克双链DNA 155.8 EB和25.9 EB的数据,超出该通道先前最高艺术密度分别达11%和52%。在映射至25°C干燥状态脱嘌呤动力学的单次编码-降解方案下,该编解码器预测以每克17.1 EB的密度可实现282年的可解码存储。这些结果使DNA存储密度接近底层通道的香农界。