Because lab accuracy of clinical speech technology systems may be overoptimistic, clinical validation is vital to demonstrate system reproducibility - in this case, the ability of the PERCEPT-R Classifier to predict clinician judgment of American English /r/ during ChainingAI motor-based speech sound disorder intervention. All five participants experienced statistically-significant improvement in untreated words following 10 sessions of combined human-ChainingAI treatment. These gains, despite a wide range of PERCEPT-human and human-human (F1-score) agreement, raise questions about best measuring classification performance for clinical speech that may be perceptually ambiguous.
翻译:由于临床言语技术系统的实验室准确性可能过于乐观,临床验证对于证明系统的可重复性至关重要——在本例中,即PERCEPT-R分类器在ChainingAI基于运动的言语障碍干预期间预测临床医生对美国英语/r/音判断的能力。所有五名参与者在经过10次人机结合ChainingAI治疗后,在未训练词汇上均表现出统计学显著改善。尽管PERCEPT-人类与人类-人类(F1分数)一致性差异范围较大,但这些增益引发了对如何最佳衡量临床言语(可能具有知觉模糊性)分类性能的疑问。