The electrocardiogram (ECG) is an accurate and widely available tool for diagnosing cardiovascular diseases. ECGs have been recorded in printed formats for decades and their digitization holds great potential for training machine learning (ML) models in algorithmic ECG diagnosis. Physical ECG archives are at risk of deterioration and scanning printed ECGs alone is insufficient, as ML models require ECG time-series data. Therefore, the digitization and conversion of paper ECG archives into time-series data is of utmost importance. Deep learning models for image processing show promise in this regard. However, the scarcity of ECG archives with reference time-series is a challenge. Data augmentation techniques utilizing \textit{digital twins} present a potential solution. We introduce a novel method for generating synthetic ECG images on standard paper-like ECG backgrounds with realistic artifacts. Distortions including handwritten text artifacts, wrinkles, creases and perspective transforms are applied to the generated images, without personally identifiable information. As a use case, we generated an ECG image dataset of 21,801 records from the 12-lead PhysioNet PTB-XL ECG time-series dataset. A deep ECG image digitization model was built and trained on the synthetic dataset, and was employed to convert the synthetic images to time-series data for evaluation. The signal-to-noise ratio (SNR) was calculated to assess the image digitization quality vs the ground truth ECG time-series. The results show an average signal recovery SNR of 27$\pm$2.8\,dB, demonstrating the significance of the proposed synthetic ECG image dataset for training deep learning models. The codebase is available as an open-access toolbox for ECG research.
翻译:心电图(ECG)是一种准确且广泛使用的心血管疾病诊断工具。数十年来,心电图多以纸质格式记录,其数字化对于训练机器学习(ML)模型实现算法化心电图诊断具有巨大潜力。由于物理心电图档案面临老化损坏的风险,且单独扫描纸质心电图尚不足以为ML模型提供所需的时间序列数据,因此将纸质心电图档案数字化并转换为时间序列数据至关重要。基于深度学习的图像处理模型在此方面展现出潜力,但缺乏伴随参考时间序列的心电图档案是一大挑战。利用数字孪生的数据增强技术提供了一种潜在解决方案。我们提出了一种新颖的方法,在标准纸质心电图背景上生成带有真实伪影的合成心电图图像。生成的图像应用了手写文本伪影、褶皱、折痕及透视变换等变形,且不包含个人身份信息。作为应用实例,我们从12导联PhysioNet PTB-XL心电图时间序列数据集中生成了21,801条记录的心电图图像数据集。基于该合成数据集构建并训练了深度心电图图像数字化模型,用于将合成图像转换为时间序列数据进行评估。通过计算信噪比(SNR)评估图像数字化质量与真实心电图时间序列的差异。结果显示平均信号恢复SNR为27±2.8 dB,证明所提出的合成心电图图像数据集对训练深度学习模型具有重要价值。该代码库作为开源工具箱供心电图研究使用。