In acupuncture therapy, the accurate location of acupoints is essential for its effectiveness. The advanced language understanding capabilities of large language models (LLMs) like Generative Pre-trained Transformers (GPT) present a significant opportunity for extracting relations related to acupoint locations from textual knowledge sources. This study aims to compare the performance of GPT with traditional deep learning models (Long Short-Term Memory (LSTM) and Bidirectional Encoder Representations from Transformers for Biomedical Text Mining (BioBERT)) in extracting acupoint-related location relations and assess the impact of pretraining and fine-tuning on GPT's performance. We utilized the World Health Organization Standard Acupuncture Point Locations in the Western Pacific Region (WHO Standard) as our corpus, which consists of descriptions of 361 acupoints. Five types of relations ('direction_of,' 'distance_of,' 'part_of,' 'near_acupoint,' and 'located_near') (n= 3,174) between acupoints were annotated. Five models were compared: BioBERT, LSTM, pre-trained GPT-3.5, and fine-tuned GPT-3.5, as well as pre-trained GPT-4. Performance metrics included micro-average exact match precision, recall, and F1 scores. Our results demonstrate that fine-tuned GPT-3.5 consistently outperformed other models in F1 scores across all relation types. Overall, it achieved the highest micro-average F1 score of 0.92. This study underscores the effectiveness of LLMs like GPT in extracting relations related to acupoint locations, with implications for accurately modeling acupuncture knowledge and promoting standard implementation in acupuncture training and practice. The findings also contribute to advancing informatics applications in traditional and complementary medicine, showcasing the potential of LLMs in natural language processing.
翻译:在针灸疗法中,穴位的准确定位对其疗效至关重要。生成式预训练变换器(GPT)等大语言模型(LLMs)所具备的先进语言理解能力,为从文本知识源中提取与穴位定位相关的关系提供了重要机遇。本研究旨在比较GPT与传统深度学习模型【长短期记忆网络(LSTM)和用于生物医学文本挖掘的双向编码器表示变换器(BioBERT)】在提取穴位相关定位关系方面的性能,并评估预训练与微调对GPT性能的影响。我们采用《世界卫生组织西太平洋区域标准穴位定位》(WHO标准)作为语料库,该语料包含361个穴位的描述。标注了五种类型的关系("方向关系"、"距离关系"、"部分关系"、"邻近穴位关系"和"位于附近关系",共计3174个关系)存在于穴位之间。对比了五种模型:BioBERT、LSTM、预训练GPT-3.5、微调GPT-3.5以及预训练GPT-4。性能指标包括微平均精确匹配精确率、召回率和F1分数。结果表明,微调后的GPT-3.5在所有关系类型的F1分数上均持续优于其他模型。总体而言,其获得了最高的微平均F1分数(0.92)。本研究凸显了GPT等大语言模型在提取穴位定位关系方面的有效性,对精准建模针灸知识及推动针灸培训与实践中的标准化实施具有指导意义。研究结果亦有助于推进传统与补充医学领域的信息学应用,展示了大语言模型在自然语言处理中的潜力。