Background: The potential of large language models (LLMs) to automate and support pharmacoepidemiologic study design is an emerging area of interest, yet their reliability remains insufficiently characterized. General-purpose LLMs often display inaccuracies, while the comparative performance of specialized biomedical LLMs in this domain remains unknown. Methods: This study evaluated general-purpose LLMs (GPT-4o and DeepSeek-R1) versus biomedically fine-tuned LLMs (QuantFactory/Bio-Medical-Llama-3-8B-GGUF and Irathernotsay/qwen2-1.5B-medical_qa-Finetune) using 46 protocols (2018-2024) from the HMA-EMA Catalogue and Sentinel System. Performance was assessed across relevance, logic of justification, and ontology-code agreement across multiple coding systems using Least-to-Most (LTM) and Active Prompting strategies. Results: GPT-4o and DeepSeek-R1 paired with LTM prompting achieved the highest relevance and logic of justification scores, with GPT-4o-LTM reaching a median relevance score of 4 in 8 of 9 questions for HMA-EMA protocols. Biomedical LLMs showed lower relevance overall and frequently generated insufficient justification. All LLMs demonstrated limited proficiency in ontology-code mapping, although LTM provided the most consistent improvements in reasoning stability. Conclusion: Off-the-shelf general-purpose LLMs currently offer superior support for pharmacoepidemiologic design compared to biomedical LLMs. Prompt strategy strongly influenced LLM performance.
翻译:背景:大型语言模型(LLMs)在自动化和支持药物流行病学研究设计方面的潜力正成为新兴研究领域,但其可靠性仍未得到充分描述。通用型LLMs常出现不准确现象,而专用生物医学LLMs在该领域的比较性能仍未知。方法:本研究评估了通用型LLMs(GPT-4o和DeepSeek-R1)与生物医学微调型LLMs(QuantFactory/Bio-Medical-Llama-3-8B-GGUF和Irathernotsay/qwen2-1.5B-medical_qa-Finetune),使用了HMA-EMA目录和哨兵系统中的46份研究方案(2018-2024年)。采用从简到繁(LTM)和主动提示策略,评估了多个编码系统下的相关性、逻辑合理性及本体编码一致性。结果:GPT-4o和DeepSeek-R1结合LTM提示在相关性和逻辑合理性评分中表现最优,其中GPT-4o-LTM在HMA-EMA方案9个问题中的8个中位相关性评分达到4。生物医学LLMs整体相关性较低,且常生成不充分的合理性论证。所有LLMs在本体编码映射方面能力有限,但LTM在推理稳定性方面提供了最一致的改进。结论:当前现成的通用型LLMs在药物流行病学设计支持方面优于生物医学LLMs。提示策略显著影响LLMs性能。