Large language models for diabetes training: a prospective study

Li, Haoxuan; Jiang, Zehua; Guan, Zhouyu; Bao, Yuqian; Liu, Yuexing; Hu, Tingting; Simó Canonge, Rafael

View

Large language models for diabetes training: a prospective study, 2025 (1.960Mb)

Author

Date

2025-03-30

Permanent link

http://hdl.handle.net/11351/12877

DOI

10.1016/j.scib.2025.01.034

ISSN

2095-9273

PMID

39947986

Show full item record

Abstract

Diabetes poses a considerable global health challenge, with varying levels of diabetes knowledge among healthcare professionals, highlighting the importance of diabetes training. Large Language Models (LLMs) provide new insights into diabetes training, but their performance in diabetes-related queries remains uncertain, especially outside the English language like Chinese. We first evaluated the performance of ten LLMs: ChatGPT-3.5, ChatGPT-4.0, Google Bard, LlaMA-7B, LlaMA2-7B, Baidu ERNIE Bot, Ali Tongyi Qianwen, MedGPT, HuatuoGPT, and Chinese LlaMA2-7B on diabetes-related queries, based on the Chinese National Certificate Examination for Primary Diabetes Care in China (NCE-CPDC) and the English Specialty Certificate Examination in Endocrinology and Diabetes of Membership of the Royal College of Physicians of the United Kingdom. Second, we assessed the training of primary care physicians (PCPs) without and with the assistance of ChatGPT-4.0 in the NCE-CPDC examination to ascertain the reliability of LLMs as medical assistants. We found that ChatGPT-4.0 outperformed other LLMs in the English examination, achieving a passing accuracy of 62.50%, which was significantly higher than that of Google Bard, LlaMA-7B, and LlaMA2-7B. For the NCE-CPFC examination, ChatGPT-4.0, Ali Tongyi Qianwen, Baidu ERNIE Bot, Google Bard, MedGPT, and ChatGPT-3.5 successfully passed, whereas LlaMA2-7B, HuatuoGPT, Chinese LLaMA2-7B, and LlaMA-7B failed. ChatGPT-4.0 (84.82%) surpassed all PCPs and assisted most PCPs in the NCE-CPDC examination (improving by 1 %–6.13%). In summary, LLMs demonstrated outstanding competence for diabetes-related questions in both the Chinese and English language, and hold great potential to assist future diabetes training for physicians globally.

Keywords

Diabetes training; Large language models; Primary diabetes care

Bibliographic citation

Li H, Jiang Z, Guan Z, Bao Y, Liu Y, Hu T, et al. Large language models for diabetes training: a prospective study. Sci Bull. 2025 Mar 30;70(6):934–42.