Purpose: Enhanced health literacy has been linked to better health outcomes; however, few interventions have been studied. We investigate whether large language models (LLMs) can serve as a medium to improve health literacy in children and other populations. Methods: We ran 288 conditions using 26 different prompts through ChatGPT-3.5, Microsoft Bing, and Google Bard. Given constraints imposed by rate limits, we tested a subset of 150 conditions through ChatGPT-4. The primary outcome measurements were the reading grade level (RGL) and word counts of output. Results: Across all models, output for basic prompts such as "Explain" and "What is (are)" were at, or exceeded, a 10th-grade RGL. When prompts were specified to explain conditions from the 1st to 12th RGL, we found that LLMs had varying abilities to tailor responses based on RGL. ChatGPT-3.5 provided responses that ranged from the 7th-grade to college freshmen RGL while ChatGPT-4 outputted responses from the 6th-grade to the college-senior RGL. Microsoft Bing provided responses from the 9th to 11th RGL while Google Bard provided responses from the 7th to 10th RGL. Discussion: ChatGPT-3.5 and ChatGPT-4 did better in achieving lower-grade level outputs. Meanwhile Bard and Bing tended to consistently produce an RGL that is at the high school level regardless of prompt. Additionally, Bard's hesitancy in providing certain outputs indicates a cautious approach towards health information. LLMs demonstrate promise in enhancing health communication, but future research should verify the accuracy and effectiveness of such tools in this context. Implications: LLMs face challenges in crafting outputs below a sixth-grade reading level. However, their capability to modify outputs above this threshold provides a potential mechanism to improve health literacy and communication in a pediatric population and beyond.
翻译:目的:健康素养的提升与更好的健康结果相关,然而相关干预措施的研究仍然较少。我们探究大语言模型(LLMs)能否作为改善儿童及其他人群健康素养的媒介。方法:我们通过ChatGPT-3.5、Microsoft Bing和Google Bard,使用26种不同提示词进行了288种条件测试。受速率限制约束,我们通过ChatGPT-4测试了其中的150种条件子集。主要结局指标为输出内容的阅读年级水平(RGL)和单词数量。结果:在所有模型中,针对“解释”和“什么是”等基础提示词的输出内容均达到或超过了10年级RGL。当提示词指定从1年级到12年级的RGL进行解释时,我们发现LLMs根据RGL定制回答的能力存在差异。ChatGPT-3.5给出的回答范围在7年级至大学新生RGL之间,而ChatGPT-4输出的回答范围在6年级至大学高年级RGL之间。Microsoft Bing提供的回答在9至11年级RGL之间,Google Bard提供的回答在7至10年级RGL之间。讨论:ChatGPT-3.5和ChatGPT-4在实现较低年级水平的输出方面表现更佳。而Bard和Bing则倾向于无论提示词如何,始终产出高中水平的RGL。此外,Bard在某些输出上的迟疑表明其对健康信息采取了谨慎态度。LLMs在增强健康沟通方面展现出潜力,但未来研究需验证此类工具在此背景下的准确性和有效性。意义:LLMs在生成低于六年级阅读水平的输出方面面临挑战。然而,它们在此阈值之上修改输出的能力,为改善儿科及其他人群的健康素养与沟通提供了潜在机制。