Automatic Speech Recognition (ASR) in medical contexts has the potential to save time, cut costs, increase report accuracy, and reduce physician burnout. However, the healthcare industry has been slower to adopt this technology, in part due to the importance of avoiding medically-relevant transcription mistakes. In this work, we present the Clinical BERTScore (CBERTScore), an ASR metric that penalizes clinically-relevant mistakes more than others. We demonstrate that this metric more closely aligns with clinician preferences on medical sentences as compared to other metrics (WER, BLUE, METEOR, etc), sometimes by wide margins. We collect a benchmark of 13 clinician preferences on 149 realistic medical sentences called the Clinician Transcript Preference benchmark (CTP), demonstrate that CBERTScore more closely matches what clinicians prefer, and release the benchmark for the community to further develop clinically-aware ASR metrics.
翻译:医学场景中的自动语音识别(ASR)技术具有节省时间、降低成本、提高报告准确性及减少医生职业倦怠的潜力。然而,医疗行业对该技术的采用速度较慢,部分原因在于避免医学相关转录错误的重要性。本研究提出临床BERTScore(CBERTScore)——一种对临床相关错误惩罚力度高于其他错误的ASR评估指标。我们证明,与其他评估指标(词错误率WER、BLEU、METEOR等)相比,该指标与临床医生对医学句子的偏好更为一致,部分情况下差距显著。我们建立了包含13位临床医生对149条真实医学句子偏好的基准测试集——临床医生转录偏好基准(CTP),验证了CBERTScore更贴近临床医生实际偏好,并公开该基准集以供业界进一步开发具有临床意识的ASR评估指标。