The online health community (OHC) is the primary channel for laypeople to share health information. To analyze the health consumer-generated content (HCGC) from the OHCs, identifying the colloquial medical expressions used by laypeople is a critical challenge. The open-access and collaborative consumer health vocabulary (OAC CHV) is the controlled vocabulary for addressing such a challenge. Nevertheless, OAC CHV is only available in English, limiting its applicability to other languages. This research proposes a cross-lingual automatic term recognition framework for extending the English CHV into a cross-lingual one. Our framework requires an English HCGC corpus and a non-English (i.e., Chinese in this study) HCGC corpus as inputs. Two monolingual word vector spaces are determined using the skip-gram algorithm so that each space encodes common word associations from laypeople within a language. Based on the isometry assumption, the framework aligns two monolingual spaces into a bilingual word vector space, where we employ cosine similarity as a metric for identifying semantically similar words across languages. The experimental results demonstrate that our framework outperforms the other two large language models in identifying CHV across languages. Our framework only requires raw HCGC corpora and a limited size of medical translations, reducing human efforts in compiling cross-lingual CHV.
翻译:在线健康社区(OHC)是非专业人士分享健康信息的主要渠道。为分析OHC中的健康消费者生成内容(HCGC),识别非专业人士使用的口语化医学术语至关重要。开放获取协作式消费者健康词汇(OAC CHV)是应对这一挑战的受控词汇表,但现有OAC CHV仅提供英文版本,限制了其在其他语言中的适用性。本研究提出一种跨语言自动术语识别框架,将英文CHV扩展为多语言版本。该框架以英文HCGC语料库和非英文(本文指中文)HCGC语料库为输入,采用skip-gram算法生成两个单语词向量空间,使每个空间编码该语言非专业人士的通用词关联模式。基于等距假设,框架将两个单语空间对齐至双语词向量空间,以余弦相似度为指标识别跨语言语义相似词。实验结果表明,本框架在识别跨语言CHV方面优于另外两个大语言模型。该框架仅需原始HCGC语料库和有限规模的医学术语翻译资源,显著减少了构建跨语言CHV所需的人工劳动。