Background: Accurate translation of radiology reports is important for multilingual research, clinical communication, and radiology education, but the validity of LLM-based evaluation remains unclear. Objective: To evaluate the educational suitability of LLM-generated Japanese translations of chest CT reports and compare radiologist assessments with LLM-as-a-judge evaluations. Methods: We analyzed 150 chest CT reports from the CT-RATE-JPN validation set. For each English report, a human-edited Japanese translation was compared with an LLM-generated translation by DeepSeek-V3.2. A board-certified radiologist and a radiology resident independently performed blinded pairwise evaluations across 4 criteria: terminology accuracy, readability, overall quality, and radiologist-style authenticity. In parallel, 3 LLM judges (DeepSeek-V3.2, Mistral Large 3, and GPT-5) evaluated the same pairs. Agreement was assessed using QWK and percentage agreement. Results: Agreement between radiologists and LLM judges was near zero (QWK=-0.04 to 0.15). Agreement between the 2 radiologists was also poor (QWK=0.01 to 0.06). Radiologist 1 rated terminology as equivalent in 59% of cases and favored the LLM translation for readability (51%) and overall quality (51%). Radiologist 2 rated readability as equivalent in 75% of cases and favored the human-edited translation for overall quality (40% vs 21%). All 3 LLM judges strongly favored the LLM translation across all criteria (70%-99%) and rated it as more radiologist-like in >93% of cases. Conclusions: LLM-generated translations were often judged natural and fluent, but the 2 radiologists differed substantially. LLM-as-a-judge showed strong preference for LLM output and negligible agreement with radiologists. For educational use of translated radiology reports, automated LLM-based evaluation alone is insufficient; expert radiologist review remains important.


翻译:背景:放射学报告的准确翻译对于多语言研究、临床沟通和放射学教育至关重要,但基于大语言模型(LLM)评估的有效性尚不明确。目的:评估LLM生成的胸部CT报告日文翻译的教育适用性,并比较放射科医生评估与LLM作为评判者的评估结果。方法:我们分析了来自CT-RATE-JPN验证集的150份胸部CT报告。针对每份英文报告,将人工编辑的日文翻译与DeepSeek-V3.2生成的LLM翻译进行比较。一名委员会认证的放射科医生和一名放射科住院医生独立进行盲审配对评估,涉及4项标准:术语准确性、可读性、整体质量和放射科医生风格的逼真度。同时,3个LLM评判者(DeepSeek-V3.2、Mistral Large 3和GPT-5)评估了相同的配对。一致性评估采用加权卡帕系数(QWK)和百分比一致性。结果:放射科医生与LLM评判者之间的一致性近乎为零(QKW=-0.04至0.15)。两名放射科医生之间的一致性也较差(QKW=0.01至0.06)。放射科医生1在59%的病例中评定术语为等效,在可读性(51%)和整体质量(51%)方面偏向LLM翻译。放射科医生2在75%的病例中评定可读性为等效,在整体质量方面偏向人工编辑翻译(40%对21%)。所有3个LLM评判者在所有标准中强烈偏向LLM翻译(70%-99%),并在超过93%的病例中评定其更接近放射科医生风格。结论:LLM生成的翻译通常被视为自然流畅,但两名放射科医生的评估存在显著差异。LLM作为评判者显示出对LLM输出的强烈偏好,并与放射科医生的一致性可忽略不计。对于翻译放射学报告的教育用途,仅依赖自动化的LLM评估是不够的;专家放射科医生的审查仍然至关重要。

0
下载
关闭预览

相关内容

《大型语言模型 (LLM) 对比研究》美海军最新报告
专知会员服务
87+阅读 · 2024年6月28日
如何检测LLM内容?UCSB等最新首篇《LLM生成内容检测》综述
LLM in Medical Domain: 大语言模型在医学领域的应用
专知会员服务
103+阅读 · 2023年6月17日
NLP 与 NLU:从语言理解到语言处理
AI研习社
15+阅读 · 2019年5月29日
中文对比英文自然语言处理NLP的区别综述
AINLP
18+阅读 · 2019年3月20日
[推荐] 这些年,我用过的点击率(CTR)预估模型!!!
菜鸟的机器学习
28+阅读 · 2017年7月31日
自然语言处理(二)机器翻译 篇 (NLP: machine translation)
DeepLearning中文论坛
12+阅读 · 2015年7月1日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
《无人机蜂群通信技术研究》50页
专知会员服务
7+阅读 · 7月19日
战力倍增器:自主武器系统与乌克兰及加沙冲突
人工智能赋能战场情报:提速决策进程
专知会员服务
6+阅读 · 7月17日
《拥抱新兴技术:面向未来军官的教育革新》
专知会员服务
8+阅读 · 7月17日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员