Purpose: The purpose of this study was to develop and evaluate rule-based algorithms to enhance the extraction of text data, including retinal nerve fiber layer (RNFL) values and other ganglion cell count (GCC) data, from Zeiss Cirrus optical coherence tomography (OCT) scan reports. Methods: DICOM files that contained encapsulated PDF reports with RNFL or Ganglion Cell in their document titles were identified from a clinical imaging repository at a single academic ophthalmic center. PDF reports were then converted into image files and processed using the PaddleOCR Python package for optical character recognition. Rule-based algorithms were designed and iteratively optimized for improved performance in extracting RNFL and GCC data. Evaluation of the algorithms was conducted through manual review of a set of RNFL and GCC reports. Results: The developed algorithms demonstrated high precision in extracting data from both RNFL and GCC scans. Precision was slightly better for the right eye in RNFL extraction (OD: 0.9803 vs. OS: 0.9046), and for the left eye in GCC extraction (OD: 0.9567 vs. OS: 0.9677). Some values presented more challenges in extraction, particularly clock hours 5 and 6 for RNFL thickness, and signal strength for GCC. Conclusions: A customized optical character recognition algorithm can identify numeric results from optical coherence scan reports with high precision. Automated processing of PDF reports can greatly reduce the time to extract OCT results on a large scale.
翻译:目的:本研究旨在开发并评估基于规则的算法,以提升从Zeiss Cirrus光学相干断层扫描(OCT)报告中提取文本数据(包括视网膜神经纤维层(RNFL)值及其他神经节细胞计数(GCC)数据)的能力。方法:从单一学术眼科中心的临床影像库中识别出文档标题包含“RNFL”或“Ganglion Cell”的封装PDF报告的DICOM文件。随后将PDF报告转换为图像文件,并使用PaddleOCR Python包进行光学字符识别。设计并迭代优化基于规则的算法,以提高RNFL及GCC数据的提取性能。通过人工复核一组RNFL及GCC报告来评估算法。结果:所开发的算法在RNFL及GCC扫描数据提取中均展现出高精度。RNFL提取中对右眼的精度略高(右眼:0.9803 vs 左眼:0.9046),而GCC提取中对左眼精度更优(右眼:0.9567 vs 左眼:0.9677)。部分数据值的提取面临更大挑战,尤以RNFL厚度中的5点和6点时钟方位,以及GCC的信号强度为甚。结论:定制的光学字符识别算法能够从光学相干扫描报告中高精度地识别数值结果。PDF报告的自动化处理可大幅缩减大规模提取OCT结果所需的时间。