切换至 "中华医学电子期刊资源库"

中华临床医师杂志(电子版) ›› 2026, Vol. 20 ›› Issue (05) : 361 -373. doi: 10.3877/cma.j.issn.1674-0785.2026.05.004

临床研究

医学专科领域大语言模型的应用与发展
许润怡1, 杨正2,3, 唐冲4, 冯珍1, 关振鹏4, 李晓5,6,()   
  1. 1 330009 南昌,南昌大学康复学院
    2 518060 深圳,香港中文大学(深圳)
    3 518172 深圳,国家健康医疗大数据研究院(深圳)工程技术研究中心
    4 100144 北京,北京大学首钢医院骨科
    5 100048 北京,中国人民解放军总医院第四医学中心骨科医学部康复医学科
    6 100048 北京,骨科与运动医学国家临床医学研究中心
  • 收稿日期:2026-05-13 出版日期:2026-05-30
  • 通信作者: 李晓
  • 基金资助:
    北京市自然科学基金-海淀原始创新联合基金资助项目(L252120)

Applications and development of large language models in specialized medicine: a systematic review

Runyi Xu1, Zheng Yang2,3, Chong Tang4, Zhen Feng1, Zhenpeng Guan4, Xiao Li5,6,()   

  1. 1 School of Rehabilitation Nanchang University, Nanchang 330009, China
    2 The Chinese University of Hong Kong, Shenzhen 518060, China
    3 The Engineering Technology Research Center, National Health Data Institute, Shenzhen 518172, China
    4 Department of Orthopedics, Peking University Shougang Hospital, Beijing 100144, China
    5 Department of Rehabilitation, Senior Department of Orthopedics, the Fourth Medical Center of PLA General Hospital, Beijing 100048, China
    6 National Clinical Research Center for Orthopedics and Sports Medicine, Beijing 100048, China
  • Received:2026-05-13 Published:2026-05-30
  • Corresponding author: Xiao Li
引用本文:

许润怡, 杨正, 唐冲, 冯珍, 关振鹏, 李晓. 医学专科领域大语言模型的应用与发展[J/OL]. 中华临床医师杂志(电子版), 2026, 20(05): 361-373.

Runyi Xu, Zheng Yang, Chong Tang, Zhen Feng, Zhenpeng Guan, Xiao Li. Applications and development of large language models in specialized medicine: a systematic review[J/OL]. Chinese Journal of Clinicians(Electronic Edition), 2026, 20(05): 361-373.

目的

以大型语言模型为代表的生成式人工智能技术,已逐步渗透到临床各医学专科的诊疗决策、科研辅助和医学教育等场景;然而,在专科适配性优化与临床落地验证等方面仍存在显著挑战。因此,本研究系统性地梳理LLMs在医学专科领域的应用,重点总结其在不同专科生成任务中的表现证据,并确定当前的挑战和未来研究方向。

方法

通过检索Pubmed、Scopus、Web of science三个核心综合性数据库,手工检索Nature、The Lancet两大世界顶级期刊,检索时间限制为2023年至2026年4月,以筛选出LLMs在医学专科领域公开可用的研究。基于纳入研究的相关数据,本研究采用描述性分析方法,对医学专科领域LLMs进行了全面的系统评价。

结果

经PubMed、Scopus、Web of Science三大核心综合性数据库检索,2023年~2026年间共25项研究符合纳入标准;另通过手工检索Nature、The Lancet 及其子刊等领域顶级期刊,补充纳入11篇符合标准的文献,最终合计纳入36项研究。纳入研究对象为医学专科LLMs,覆盖内科、外科、眼科、肿瘤科、中医全科等8类专科,以单模态模型为主(75%),多模态模型仅占25%;应用以临床诊疗辅助与医学教育为主,但普遍存在专科适配不足、多模态数据利用不充分、临床转化证据缺乏等问题。

结论

大型语言模型(LLMs)在医学专科领域已展现出显著的应用潜力,但目前仍普遍存在专科适配不足、多模态数据利用不充分、临床转化证据缺乏等问题。未来需聚焦这一核心问题深度开发医学专科LLMs,构建以临床需求为导向、由临床医生主导、多模态融合的统一验证与评估标准框架,形成幻觉评估、可解释性、伦理监管协同的三位一体规范体系,并开展进一步的研究。

Objective

Generative artificial intelligence (AI) technologies, represented by large language models (LLMs), have gradually permeated various clinical scenarios across medical specialties, including diagnostic and treatment decision-making, research support, and medical education; however, significant challenges remain in terms of optimizing specialty-specific adaptation and validating clinical implementation. Therefore, this study systematically reviews the applications of LLMs in medical specialties, focusing on summarizing performance evidence across various specialty-specific generation tasks, and identifies current challenges and future research directions.

Methods

By searching the three core comprehensive databases-PubMed, Scopus, and Web of Science-and manually searching the two world-leading journals, Nature and The Lancet, with a search period limited to 2023 through April 2026, we identified publicly available studies on LLMs in medical specialties. Based on data from the included studies, this study employed a descriptive analysis method to conduct a comprehensive systematic review of LLMs in medical specialties.

Results

Searches of the three core comprehensive databases identified 25 studies meeting the inclusion criteria. Additionally, a manual search of top-tier journals in the field, including Nature, The Lancet, and their subsidiary journals, yielded 11 additional studies meeting the criteria, resulting in a total of 36 included studies. The included studies focused on medical specialty LLMs, covering eight specialties including internal medicine, surgery, ophthalmology, oncology, and traditional Chinese medicine general practice. Monomodal models constituted the majority (75%), while multimodal models accounted for only 25%. Applications were primarily concentrated in clinical diagnosis, treatment assistance, and medical education; however, common issues included insufficient specialty adaptation, inadequate utilization of multimodal data, and a lack of evidence for clinical translation.

Conclusion

LLMs have demonstrated significant application potential in medical specialties. However, issues such as insufficient specialty adaptation, inadequate utilization of multimodal data, and a lack of clinical translation evidence remain prevalent. Future efforts should focus on addressing these core issues by deeply developing specialty-specific LLMs, establishing a unified validation and evaluation framework that is clinically driven, clinician-led, and multimodal, and creating a three-pronged regulatory system that integrates hallucination assessment, explainability, and ethical oversight.

图1 文献筛选流程
表1a 模型基本信息
模型 领域 发表期刊 开发机构 时间 模型类型
RETFound[10] 眼科 Nature 伦敦大学等 2023 单模态大模型
EyeGPT[11] 眼科 JMIR 中国香港理工大学 2024 单模态大模型
MOPH[12] 眼科 BJO 上海交大新华医院等 2024 单模态大模型
TCMchat[13] 中医全科 Pharmacol.Res 浙江大学等 2024 单模态大模型
MedChatZH[14] 中医全科 Comput Biol Med 华东理工大学等 2024 单模态大模型
TCMLLM-PR[15] 中医全科 Digit Chin Med 北京交通大学等 2024 单模态大模型
ChatENT[16] 耳鼻喉头颈外科 Otolaryngol Head Neck Surg 阿尔伯塔大学耳鼻喉头颈外科等 2024 单模态大模型
PneumoLLM[17] 呼吸内科 Med Image Anal 中国医学科学院等 2024 单模态模型
PlasticSurgeryGPT[18] 整形外科 Aesthet Surg J 克利夫兰诊所等 2024 单模态大模型
DeepDR-LLM[7] 内分泌科 Nat.Med 上海交通大学医学院附属第六人民医院等 2024 多模态大模型
Ophtimus-V2-Tx[19] 眼科 Sci.Rep 韩国庆尚国立大学等 2025 单模态大模型
DeLLiriuM[20] 重症医学科 Sci.Rep 佛罗里达大学等 2025 单模态大模型
NeonatalBERT[21] 儿科 Lancet Digit. Health 斯坦福大学等 2025 单模态大模型
MenstLLaMA[22] 妇科 JMIR 印度理工学院德里分校 2025 单模态大模型
Woollie[23] 肿瘤科 NPJ Digit Med 美国纪念斯隆凯特琳癌症中心等 2025 单模态大模型
GutGPT[24] 消化科 Biomed.Inform 山西省人民医院等 2025 单模态大模型
EYE-Llama[25] 眼科 iScience 美国北卡大学夏洛特分校等 2025 单模态大模型
Hypnos[26] 麻醉科 Neurocomputing 中国石油大学等 2025 单模态大模型
MedicalGLM[27] 儿科 J Biomed Inform 山东中医药大学等 2025 单模态大模型
TCM-KLLaMA[28] 中医全科 Comput Biol Med 浙江工商大学等 2025 单模态大模型
Qibo[29] 中医全科 Expert Syst Appl 天津中医药大学等 2025 单模态大模型
TCMLCM[30] 中医全科 Digit Chin Med 南京中医药大学等 2025 单模态大模型
BianCang[31] 中医全科 IEEE J Biomed Health Inform 山东中医药大学附属医院等 2025 单模态大模型
XuanHuGPT[32] 中医全科 Chin Med 澳门理工大学等 2025 单模态大模型
ZhongdaChat-ED[33] 男科 Asian J Androl 中国东南大学等 2025 单模态大模型
DeepGEM[34] 肿瘤科 Lancet Oncol 广州医科大学附属第一医院等 2025 多模态大模型
HepaPathGPT[8] 肿瘤科 EBioMedicine 清华大学等 2025 多模态大模型
EyeFM[35] 眼科 Nat.Med 上海交通大学等 2025 多模态大模型
NeoCLIP[36] 儿科 NPJ Digit Med 美国贝斯以色列女执事医疗中心等 2025 多模态大模型
RadFM[37] 医学影像科 Nat.Commun. 中国上海交通大学等 2025 多模态大模型
CXR-LLaVA[38] 医学影像科 EUR RADIOL 韩国首尔国立大学医院等 2025 多模态大模型
GraphRAG-GDM[39] 内分泌科 JMIR Diabetes 扎耶德大学等 2026 单模态大模型
HerbWise[40] 中医全科 Chin Herb Med 成都中医药大学等 2026 单模态大模型
CancerLLM[41] 肿瘤科 NPJ Digit Med 明尼苏达大学等 2026 单模态大模型
OBUSight[42] 眼科 Adv.Sci 浙江大学、中山大学等 2026 多模态大模型
SpiroLLM[9] 呼吸与危重症医学科 PLOS Digit Health 西安电子科技大学等 2026 多模态大模型
表1b 模型输入数据与任务
模型 输入数据 任务指令
RETFound[10] 视网膜图像 眼病检测、系统性疾病预测
EyeGPT[11] 临床文本+专业文献+知识库 眼科知识问答、患者咨询
MOPH[12] 眼科教材、指南、考试题库、门诊病例 眼科答题、循证问答、病例诊断
TCMchat[13] 中医药知识库+临床文本 中医知识问答、方剂推荐
MedChatZH[14] 中医古籍、医疗指令、临床问诊数据 中医辨证、方药、调理、穴位
TCMLLM-PR[15] 中医教材、药典、肺/肝/中风/糖尿病/脾胃病历 多疾病中医处方推荐、症状-方药匹配
ChatENT[16] 耳鼻喉专科知识库、指南、文献 专科问答、考试答案、决策建议
PneumoLLM[17] 胸部X光影像、肺分割数据、临床标签 尘肺病诊断、病变判别、影像报告
PlasticSurgeryGPT[18] 2.5万篇整形外科文献摘要 整形知识、文献解读、专业问答
DeepDR-LLM[7] 糖尿病患者临床文本数据+眼底图像 DR分级和筛查、个性化糖尿病管理建议
Ophtimus-V2-Tx[19] 眼科病例报告文本 生成主要诊断、手术方案、用药
DeLLiriuM[20] ICU结构化电子病历 ICU谵妄风险预测
NeonatalBERT[21] 新生儿临床记录(入院、病程记录等) 新生儿疾病风险预测
MenstLLaMA[22] 月经、妇科临床文本 妇科健康咨询、月经问题解答
Woollie[23] 肿瘤放射报告文本 肿瘤进展预测、临床决策支持
GutGPT[24] 胃肠病学临床指南、病例文本 胃肠疾病诊断、治疗建议
EYE-Llama[25] 眼科文献、教材、临床文本 眼科问答、诊断、治疗、考试答题
Hypnos[26] 通用医疗QA、麻醉真实或合成数据 麻醉知识、方案、用药、病例分析
MedicalGLM[27] 儿科知识数据集、临床文本 儿科问答、饮食、治疗、用药
TCM-KLLaMA[28] 中医症状、舌诊脉诊知识图谱、症状-处方对 中医辨证、标准化中药处方、药材推荐
Qibo[29] 中医古籍、教材、方剂、知识图谱、临床文本 中医辨证、方剂推荐、多轮问诊、知识问答
TCMLCM[30] 中医肺癌知识图谱、临床文本、临床指南 中医肺癌问答、辨证结论、方药推荐
BianCang[31] 中医典籍、药典、四诊信息、真实临床病历 中医辨证分型、疾病诊断、治法处方
XuanHuGPT[32] 10万条中医结构化数据(古籍、方剂、药理) 中医问答、辨证推理、方药解释
ZhongdaChat-ED[33] ED指南、健康咨询、临床文本 ED个性化咨询、临床诊疗方案
DeepGEM[34] 肺癌患者组织病理学全切片图像+临床文本数据 预测多种肺癌驱动基因突变、生成基因突变空间分布图
HepaPathGPT[8] 肝癌增强CT/MRI影像+病理报告 术前病理特征预测生成类病理报告
EyeFM[35] OCT/FFF等眼科影像+临床文本 眼科多模态诊断分诊
NeoCLIP[36] 新生儿X光片+放射报告 识别新生儿X光片病理特征、医疗设备
RadFM[37] 多模态放射学数据+临床文本 放射诊断、报告生成、医学VQA
CXR-LLaVA[38] 胸部X光图像+放射报告 生成放射学报告、检测主要病理发现
GraphRAG-GDM[39] GDM论文、指南、诊疗数据 GDM风险评估、方案推荐、解释
HerbWise[40] 草药、基因组、化学成分 草药问答、化合物预测、实体抽取
CancerLLM[41] 癌症病历、临床笔记、病理报告 肿瘤表型抽取、癌症诊断生成
OBUSight[42] 眼部B超图像+临床报告 眼科超声解读、疾病诊断预测
SpiroLLM[9] 肺功能时间序列曲线、PFT数值、人口学信息 COPD诊断报告、肺功能解读、曲线形态描述
表1c 模型应用场景
模型 应用场景
RETFound[10] 眼底病筛查、全身健康预警
EyeGPT[11] 眼科科普、医患沟通
MOPH[12] 眼科规培、基层诊断、患者教育
TCMchat[13] 中医辅助问诊
MedChatZH[14] 中医问诊、科普、辅助诊疗、教学
TCMLLM-PR[15] 中医处方智能推荐、跨病种迁移
ChatENT[16] 教学、患者教育、规培考试、临床辅助
PneumoLLM[17] 职业肺病筛查、AI辅助影像诊断
PlasticSurgeryGPT[18] 整形教学、临床咨询、研究辅助
DeepDR-LLM[7] 糖尿病眼底病与慢病管理
Ophtimus-V2-Tx[19] 眼科临床诊断、治疗规划
DeLLiriuM[20] 重症监护早期预警
NeonatalBERT[21] 新生儿重症监护预后评估
MenstLLaMA[22] 月经健康教育与咨询
Woollie[23] 肿瘤诊疗辅助
GutGPT[24] 消化内科辅助诊疗
EYE-Llama[25] 眼科教学、患者教育、临床辅助
Hypnos[26] 麻醉培训、临床辅助、围手术期管理
MedicalGLM[27] 儿科咨询、医师培训、家长指导
TCM-KLLaMA[28] 中医临床辅助开方、处方推荐
Qibo[29] 中医智能问诊、临床辅助、中医药教学
TCMLCM[30] 中医肺癌辅助诊疗、专业问答
BianCang[31] 中医临床辨证、中西医结合诊断
XuanHuGPT[32] 中医科普、教学、辨证支持
ZhongdaChat-ED[33] 医师辅助决策、隐私咨询
DeepGEM[34] 肿瘤学精准医疗
HepaPathGPT[8] 肝胆肿瘤学术前评估
EyeFM[35] 眼科全流程临床辅助决策
NeoCLIP[36] 新生儿重症监护
RadFM[37] 通用放射科辅助诊疗
CXR-LLaVA[38] 胸部影像辅助诊断
GraphRAG-GDM[39] 妊娠糖尿病辅助诊疗、患者科普
HerbWise[40] 中药研究、新药发现、质量控制
CancerLLM[41] 肿瘤信息提取、病理结构化、辅助诊断
OBUSight[42] 眼科超声影像辅助诊断
SpiroLLM[9] 慢阻肺筛查、肺功能报告生成
图2 输入数据类型分布图 注:部分模型可适配多种输入数据类型,因此各分类占比之和大于100%
图3 应用场景类型分布图 注:部分模型可适配多种应用场景类型,因此各分类占比之和>100%
图4 医学大模型构建路径图
表2a 模型基础特征与人工评估指标
模型 基础模型 可用性 人工评估指标
RETFound[10] 基座:MAE
主干:ViT-L/16
开源 NR
EyeGPT[11] Llama 2-7B-chat 开源 NR
MOPH[12] ChatGLM2-6B 闭源 指南依从率:83.3%;差/极差回答(1–:6.7%;潜在误导回答:10%;专家评分一致性ICC:0.91、0.95
TCMchat[13] Baichuan2-7B-Chat 开源 NR
MedChatZH[14] Baichuan-7B 开源 NR
TCMLLM-PR[15] ChatGLM-6B 开源 NR
ChatENT[16] ChatGPT4.0 闭源 NR
PneumoLLM[17] LLaMA-7B 开源 NR
PlasticSurgeryGPT[18] GPT-2 开源 NR
DeepDR-LLM[7] LLaMA-7B 开源 NR
Ophtimus-V2-Tx[19] LLaMA3.18B 开源 NR
DeLLiriuM[20] GatorTronS 开源 NR
NeonatalBERT[21] Bio-ClinicalBERT 开源 NR
MenstLLaMA[22] Meta-Llama-3-8B-Instruct 开源 临床专家评分相关性:3.97/5;可理解性:4.48/5;精准度:3.90/5;正确性:4.00/5;上下文敏感性:3.41/5;执业医师评分相关性:3.5/5;可理解性:3.6/5;精准度:3.1/5;正确性:3.5/5;上下文敏感性:4.0/5;普通用户评分可理解性:4.7/5;相关性:4.3/5;精准度:4.28/5;正确性:4.1/5;上下文敏感性:3.9
Woollie[23] NR 闭源 NR
GutGPT[24] Baichuan-13B-Chat 闭源 专家评估(200道消化内科题)平均准确率得分:9.31/10;“优秀”评级占比:73.5%;诊断准确率较基线提升:9.59%;人文关怀达标率:100%
EYE-Llama[25] Llama 2-7B-Chat 开源 眼科专家评分总分:11.6/20;一致性:14;危害性:10.2
Hypnos[26] ChatGLM3-6B 闭源 NR
MedicalGLM[27] ChatGLM-6B 开源 NR
TCM-KLLaMA[28] Chinese-LLaMA2-7B 闭源 NR
Qibo[29] Chinese-LLaMA-7B / 13B 闭源 安全性(胜率:38%~95%;平局率:3%~29%;负率:1%~33%);专业性(胜率39%~96%;平局率1%~33%;负率3%~38%);流利
度(胜率32%~96%;平局率:2%~37%;负率:2%~35%)
TCMLCM[30] ChatGLM2-6B 闭源 中医肿瘤专家评分准确性:52%;专业性:46%;可用性:48%
BianCang[31] Qwen-2/Qwen2.5(7B/14B) 开源 专家主观胜率,专业性胜率:73%;流畅性胜率:65%;安全性胜率:85%;
XuanHuGPT[32] ChatGLM2-6B 开源 中医专家盲评:专业性:41%~80%;安全性:25%
ZhongdaChat-ED[33] DeepSeek-R1-32B 闭源 消费者版(患者咨询)准确性:4.77/5;人文关怀:4.86/5;易懂性:4.88/5;专业版(临床决策)临床意义得分≥85.2%;知识前沿性4.52/5;临床有效响应率:100%
DeepGEM[34] CTransPath 开源 NR
HepaPathGPT[8] LLaVA 1.5-7B 开源 5名病理医生评分,准确性接受率:92.5%;完整性接受率:87.4%;逻辑一致性、专业性、实用性均>85%
EyeFM[35] 视觉模块:MultiMAE
语言模块:LLaMA 2
开源 眼科医师VQA评分:15分(IQR 12~15);同理心评分:5/5;报告标准化评分:干预组37,对照组33;患者自我管理依从性:70.1%;医生系统可用性评分:92.5/100
NeoCLIP[36] 视觉模块:ResNet-50
文本模块:BioViL-T
闭源 NR
RadFM[37] 视觉模块:3D ViT
文本模块:MedLLaMA-13B
开源 放射科专家评分5分制,医学VQA得分:2.87;报告生成得分:1.88;原理诊断得分:1.76;平均分:2.17
CXR-LLaVA[38] ViT-L/16 +LLaMA 2-7B 开源 放射科医生主观评估,无需修改可直接使用:51.3%;自主报告成功率:72.7%;真实报告成功率:84.0%
GraphRAG-GDM[39] 本地部署大语言模型 开源 相关性评分:接近1.0(满分)
HerbWise[40] ChatGLM3-6B 闭源 NR
CancerLLM[41] Mistral-7B 开源 NR
OBUSight[42] 视觉模块:ResNet-101
文本模块:BERT
闭源 评分者一致性,科恩kappa系数=0.923;可直接使用/小幅修改报告占比:74.5%;完全废弃报告:10.0%
SpiroLLM[9] 时序模块:SpiroEncoder
文本模块:Llama3.1-8B
对齐模块:SpiroProjector(MLP)
开源 呼吸科专家评分:
事实准确性:78.36;完整性:86.39逻辑合理性:81.63;医学术语:95.62医学安全性:89.03;曲线描述:85.76
表2b 模型自动评估指标
模型 自动评估指标
RETFound[10] 糖尿病视网膜病AUROC:0.943;青光眼分类AUROC:0.90+;眼病预后预测,对侧眼1年湿性AMD转(CFP):0.862;对侧眼1年湿性AMD转化(OCT):0.799;全身疾病预测(3年发病)心肌梗死(MI):0.737;心力衰竭:0.794;缺血性卒中:0.754;帕金森病:0.669
EyeGPT[11] 原始Llama2幻觉率:80.8%;最佳微调模型幻觉率:44.2%;最佳微调+教科书检索幻觉率:40.8%
MOPH[12] 眼科考试总分:64.7;单选题正确率:56.0%;问答题得分率:73.4%;临床诊断准确率:81.1%;白内障亚专科准确率:96.7%;视网膜疾病准确率:63.3%;眼科专科题准确率:57.4%
TCMchat[13] 选择题:草药知识准确率:71.6%;方剂知识准确率:76.8%;阅读理解:BLEU:0.584;METEOR:0.737;ROUGE-L:0.766;病例诊断准确率:0.847;精确率:0.52;召回率:0.592;F1-score:0.5784;实体抽取精确率:0.975;召回率:0.861;F1-score:0.907
MedChatZH[14] BLEU-1:56.14;BLEU-4:9.17;ROUGE-L:21.77;GLEU:10.32;
TCMLLM-PR[15] 中医教材数据集AUROC:0.5914;CHP AUROC药典数据集:0.2818;跨数据集(教材→肝病)AUROC:0.1551;
ChatENT[16] 开放式简答题(加拿大皇家学院考试)有效性:95.7%;胜任力得分:87.2%;错误率相对下降:58.4%(对比GPT-4.0);多选题(美国BoardVitals题库)总体正确率:80%;GPT-4.0正确率:73%;错误率相对下降:26%
PneumoLLM[17] 准确率:75.87%;灵敏度:80.54%;特异度:67.66%;AUC:78.98%
PlasticSurgeryGPT[18] BLEU:0.1355;METEOR:0.5836;ROUGE-1:0.2168
DeepDR-LLM[7] 标准眼底图AUROC:0.892–0.933;便携眼底图AUROC:0.896~0.920
Ophtimus-V2-Tx[19] BLEU:0.26;ROUGE-L:0.40;METEOR:0.45;SEMSCORE:0.80
DeLLiriuM[20] 外部验证集(MIMIC+eICU):AUROC:82.4%;AUPRC:11.8%);内部验证集(UFH):AUROC:84.8%;AUPRC:19.2%
NeonatalBERT[21] 平均AUROC(主队列):0.910;平均AUPRC(主队列):0.291;平均F1-score:0.416;外部队列平均AUROC:0.935;外部队列平均AUPRC:0.360
MenstLLaMA[22] BLEU:0.059;ROUGE-L:0.224;METEOR:0.290;BERTScore:0.911
Woollie[23] 内部验证AUROC:0.97;外部验证 AUROC:0.88
GutGPT[24] CMB数据集准确率:67.53%;CMExam数据集准确率:71.84%;平均准确率:69.69%
EYE-Llama[25] PubMedQA准确率:0.96;PubMedQA F1:0.74;MedMCQA准确率:0.39;MedMCQA F1:0.23;开放式对话BERTScore F1:0.5715;BARTScore:0.7035;BLEU:0.0227
Hypnos[26] 多选题准确率:中医知识 0.826;草药基因组 0.720;实体抽取 F1:草药实体 0.9057;化合物实体0.7385;BLEU-4:THM-QA 0.1765;CC-QA 0.3908;HG-QA 0.1700;ROUGE-L:THM-QA 0.4172;CC-QA 0.4957;HG-QA 0.5547;ADMET预测ACC:73.85%;AUC:0.7580
MedicalGLM[27] ROUGE-L:44.50;BLEU:32.61;ROUGE-1:54.90;ROUGE-2:28.02
TCM-KLLaMA[28] 精确率:0.229;召回率:0.310;F1-Score:0.263
Qibo[29] ROUGE-L:TCM-NER 0.72、TCM-RC 0.61、TCM-SD 0.54;BLEU-1:TCM-NER 0.72、TCM-RC 0.70、TCM-SD 0.50
TCMLCM[30] 准确率:79.68%;BLEU:32.15%;ROUGE-L:59.08%;TCM-LCEval 平均分:38.30%
BianCang[31] TCM辨证准确率:最高82.10%;TCM疾病诊断准确率:最高89.40%;MLEC-TCM准确率:最高92.39%;CMB 医学基准准确率:最高84.34%
XuanHuGPT[32] BLEU-1:51.53;ROUGE-L:32.24;METEOR:0.260
ZhongdaChat-ED[33] NR
DeepGEM[34] EGFR:AUC:0.96;准确率:0.95;F1-Score:0.95;精确率:0.96;召回率:0.95;KRAS:AUC:0.82;准确率:0.79;F1-Score:0.80;精确率:0.83;召回率:0.79;TP53:AUC:0.89;准确率:0.86;F1-Score:0.86;精确率:0.86;召回率:0.86;LRP1B:AUC:0.86;准确率:0.80;F1-Score:0.80;精确率:0.81;召回率:0.80;Median(IQR):AUC:0.87;准确率:0.83;F1-Score:0.83;精确率:0.85;召回率:0.83
HepaPathGPT[8] 肿瘤分割指标mIoU:0.883±0.007;Dice系数:0.934±0.006;病理报告生成指标;BLEU-4:62.7;ROUGE-1:84.2;ROUGE-2:71.7;ROUGE-L:73.6;病理标志物预测(外部验证,6项指标平均)准确率:0.697
EyeFM[35] 疾病检测:AUROC(最高0.932)病灶分割:Dice系数;跨模态诊断:AUROC 0.883–0.907;RCT主要指标;正确诊断率:92.2%;正确转诊率:92.2%
NeoCLIP[36] 支气管肺发育不良AUROC:0.94;气管插管定位AUROC:0.97;脐静脉导管定位AUROC:0.97;气胸AUROC:0.93;LLM标签提取准确率:0.90;灵敏度:0.98;特异度:0.84
RadFM[37] 诊断任务准确率:72.95%;F1-Score:92.21%;报告生成/VQABLEU-1:83.16;ROUGE-L:83.65
CXR-LLaVA[38] MIMIC内部测试集:准确度:86%;精确度:89%;F1-score:81%;CheXpert内部测试集:准确度:74%;精确度:68%;F1-score:67%外部测试集:准确度:90%;精确度:93%;F1-score:56%
GraphRAG-GDM[39] BLEU:0.99;Jaccard 相似度:0.98;BERTScore:0.98
HerbWise[40] 多选题准确率:中医知识82.6%、草药基因组72.0%;实体抽取F1-Score:草药实体0.9057、化合物实体0.7385;BLEU-4:THM-QA 0.1765、CC-QA 0.3908、HG-QA 0.1700;ROUGE-L:THM-QA 0.4172、CC-QA 0.4957、HG-QA 0.5547;BLEU-1:基因组理解0.5544、中医文本理解0.5880
CancerLLM[41] 癌症诊断生成:平均F1:86.81%;Exact Match F1:83.50%;BLEU-2 F1:86.60%;ROUGE-L F1:90.34%;独立患者列队验证:平均F1:85.08%;Exact Match F1:82.10%;BLEU-2 F1:82.10%;ROUGE-L F1:91.05%;癌症表征提取:平均F1:91.78%;Exact Match F1:89.37%;BLEU-2 F1:91.98%;ROUGE-L F1:93.98%
OBUSight[42] 准确率:77.55;F1-Score:75.99;BLEU-1:0.867;BLEU-2:0.834;BLEU-3:0.806;BLEU-4:0.783;ROUGE-L:0.859;METEOR:0.577
SpiroLLM[9] COPD诊断F1-Score:0.8947;AUROC:0.8977
1
关于促进和规范“人工智能+医疗卫生”应用发展的实施意见 [J]. 中华人民共和国国家卫生健康委员会公报, 2025(12): 33-36.
2
中共中央关于制定国民经济和社会发展第十五个五年规划的建议 [J]. 工业信息安全, 2025 (6): 82-96.
3
Vora LK, Gholap AD, Jetha K, et al. Artificial intelligence in pharmaceutical technology and drug delivery design [J]. Pharmaceutics, 2023, 15(7): 1916.
4
Reddy S. Generative AI in healthcare:an implementation science informed translational path on application,integration and governance [J]. Implement Sci, 2024, 19(1): 27-41.
5
Cho YS. From code to cure:unleashing the power of generative artificial intelligence in medicine [J]. Int Neurourol J, 2023, 27(4): 225-226.
6
Foote HP, Hong C, Anwar M,et al. Embracing generative artificial intelligence in clinical research and beyond [J]. JACC Adv, 2025, 4(3): 101593.
7
Li J, Guan Z, Wang J, et al. Integrated image-based deep learning and language models for primary diabetes care [J]. Nat Med, 2024, 30(10): 2886-2896.
8
Wang L, Tian F, Li F, et al. A generative vision-language model for holistic pathological assessment using preoperative imaging in hepatocellular carcinoma [J]. EBio Medicine, 2025, 122: 106060.
9
Mei S, Long Y, Xiao X, et al. SpiroLLM: finetuning pretrained LLMs to understand spirogram time series with clinical validation in COPD reporting [J]. PLOS Digit Health, 2026, 5(3): e0001300.
10
Zhou Y, Chia MA, Wagner SK, et al. A foundation model for generalizable disease detection from retinal images [J]. Nature, 2023, 622(7981): 156-163.
11
Chen X, Zhao Z, Zhang W, et al. EyeGPT for patient inquiries and medical education: development and validation of an ophthalmology large language model [J]. J Med Internet Res, 2024, 26: e60063.
12
Zheng C, Ye H, Guo J, et al. Development and evaluation of a large language model of ophthalmology in Chinese [J]. Br J Ophthalmol, 2024, 108(10): 1390-1397.
13
Dai Y, Shao X, Zhang J, et al. TCMChat: a generative large language model for traditional Chinese medicine [J]. Pharmacol Res, 2024, 210: 107530.
14
Tan Y, Zhang Z, Li M, et al. MedChatZH: a tuning LLM for traditional Chinese medicine consultations [J]. Comput Biol Med, 2024, 172: 108290.
15
Haoyu T, Kuo Y, Xin D, et al. TCMLLM-PR: evaluation of large language models for prescription recommendation in traditional chinese medicine [J]. Digit Chin Med, 2024, 7(4): 343-355.
16
Long C, Subburam D, Lowe K, et al. ChatENT: augmented large language model for expert knowledge retrieval in otolaryngology-head and neck surgery [J]. Otolaryngol Head Neck Surg, 2024, 171(4): 1042-1051.
17
Song M, Wang J, Yu Z, et al. PneumoLLM: Harnessing the power of large language model for pneumoconiosis diagnosis [J]. Med Image Anal, 2024, 97: 103248.
18
Ozmen BB. Initial proof-of-concept study for a plastic surgery–specific artificial intelligence large language model: PlasticSurgeryGPT [J]. Aesthet Surg J, 2025, 45(8): 860-864.
19
Kwon M, Jang KJ, Baek SJ, et al. Ophtimus-V2-tx: a compact domain-specific LLM for ophthalmic diagnosis and treatment planning [J]. Sci Rep, 2025, 15(1): 43532.
20
Contreras M, Kapoor S, Zhang J, et al. A large language model for delirium prediction in the intensive care unit using structured electronic health records [J]. Sci Rep, 2025, 15(1): 38890.
21
Xie F, Chung P, Reiss JD, et al. Development and validation of a pre-trained language model for neonatal morbidities: a retrospective, multicentre, prognostic study [J]. Lancet Digit Health, 2025, 7(12): 100926.
22
Adhikary PK, Motiyani I, Oke G, et al. Menstrual health education using a specialized large language model in India: development and evaluation study of MenstLLaMA [J]. J Med Internet Res, 2025, 27: e71977.
23
Heydari K, Enichen EJ, Li B, et al. Incorporating large language models as clinical decision support in oncology: the Woollie model [J]. NPJ Digit Med, 2025, 8(1): 529.
24
Zhang RY, Qiang PP, Hao YX, et al. GutGPT: a multidimensional knowledge-enhanced large language model for gastrointestinal medicine [J]. J Biomed Inform, 2025, 169: 104885.
25
Haghighi T, Gholami S, Sokol JT, et al. EYE-Llama, an in-domain large language model for ophthalmology [J]. iScience, 2025, 28(7): 112984.
26
Wang Z, Jiang J, Zhan Y, et al. Hypnos: a domain-specific large language model for anesthesiology [J]. IJON, 2025, 624: 129389.
27
Wang X, Sun Z, Wang P, et al. MedicalGLM: a pediatric medical question answering model with a quality evaluation mechanism [J]. J Biomed Inform, 2025, 165: 104793.
28
Zhuang Y, Yu L, Jiang N, et al. TCM-KLLaMA: intelligent generation model for traditional Chinese medicine prescriptions based on knowledge graph and large language model [J]. Comput Biol Med, 2025, 189: 109887.
29
Jia Y, Ji X, Wang X, et al. Qibo: a large language model for traditional Chinese medicine [J]. Expert Syst Appl, 2025, 284: 127672.
30
Chunfang Z, Qingyue G, Wendong Z, et al. TCMLCM: an intelligent question-answering model for traditional Chinese medicine lung cancer based on the KG2TRAG method [J]. Digit Chin Med, 2025, 8(1): 36-45.
31
Wei S, Peng X, Wang Y, et al. BianCang: a traditional Chinese medicine large language model [J]. IEEE J Biomed Health Inform, 2025.
32
Tong X, Ding X, Jia H, et al. XuanHuGPT: parameter-efficient fine-tuning of large language model in the field of traditional Chinese medicine [J].Chin Med, 2025, 20(1): 204-229.
33
Xia Y, Zhu YK, Liu CH, et al. ZhongdaChat-ED: a medical large language model for personalized erectile dysfunction health consultation and professional clinical decision-making using retrieval-augmented generation [J]. Asian J Androl, 2026, 28(1): 71-79.
34
Zhao Y, Xiong S, Ren Q, et al. Deep learning using histological images for gene mutation prediction in lung cancer: a multicentre retrospective study [J]. Lancet Oncol, 2025, 26(1): 136-146.
35
Wu Y, Qian B, Li T, et al. An eyecare foundation model for clinical assistance: a randomized controlled trial [J]. Nat Med, 2025, 31(10): 3404-3413.
36
Huang Y, Sharma P, Palepu A, et al. NeoCLIP: a self-supervised foundation model for the interpretation of neonatal radiographs [J]. NPJ Digit Med, 2025, 8(1): 570.
37
Wu C, Zhang X, Zhang Y, et al. Towards generalist foundation model for radiology by leveraging web-scale 2D&3D medical data [J]. Nat Commun, 2025, 16(1): 7866.
38
Lee S, Youn J, Kim H, et al. CXR-LLaVA: a multimodal large language model for interpreting chest X-ray images [J]. Eur Radiol, 2025, 35(7): 4374-4386.
39
Evangelista E, Ruba F, Bukhari S, et al. GraphRAG-enabled local large language model for gestational diabetes mellitus: development of a proof-of-concept [J]. JMIR Diabetes, 2026, 11: e76454.
40
Xiao W, Cheng Q, Mao Y, et al. HerbWise: a domain-specific large language model for traditional herbal medicine [J]. Chin Herb Med, 2026: S1674638426000122.
41
Li M, Zhan Z, Huang J, et al. CancerLLM: a large language model in cancer domain [J]. NPJ Digit Med, 2026, 9(1): 266-287.
42
Liu X, Shao A, Guan B, et al. OBUSight: clinically aligned generative AI for ophthalmic ultrasound interpretation and diagnosis [J]. Adv Sci, 2026, 13(16): e15864.
43
Thirunavukarasu AJ, Ting DSJ, Elangovan K, et al. Large language models in medicine [J]. Nat Med, 2023, 29(8): 1930-1940.
44
Bergquist M, Rolandsson B, Gryska E, et al. Trust and stakeholder perspectives on the implementation of AI tools in clinical radiology [J]. Eur Radiol, 2023, 34(1): 338-347.
45
Wagner MM, Hogan WR, Levander JD, et al. Towards machine-FAIR: representing software and datasets to facilitate reuse and scientific discovery by machines [J]. J Biomed Inform, 2024, 154: 104647.
46
Busch F, Hoffmann L, Rueger C, et al. Current applications and challenges in large language models for patient care: a systematic review [J]. Commun Med, 2025, 5(1): 26.
47
Li H, Moon JT, Purkayastha S, et al. Ethics of large language models in medicine and medical research [J]. Lancet Digit Health, 2023, 5(6): e333-e335.
48
Reverberi C, Rigon T, Solari A, et al. Experimental evidence of effective human-AI collaboration in medical decision-making [J]. Sci Rep, 2022, 12(1): 14952.
49
Zhao L, Liu S, Xin T, et al. AI agent in healthcare: applications, evaluations, and future directions [J]. NPJ Digit Med, 2026, 2(1): 31-42.
50
李小鹏, 喻慧心, 常青, 等. 医疗基础大模型的应用现状、挑战及展望 [J]. 医学信息学杂志, 2026, 47(1): 16-23.
51
Wang S, Tang Z, Yang H, et al. A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains [J]. NPJ Digit Med, 2025, 9(1): 91-105.
52
Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge [J]. Nat, 2023, 620(7972): 172-180.
53
Ong JCL, Chang SY, William W, et al. Ethical and regulatory challenges of large language models in medicine [J]. Lancet Digit Health, 2024, 6(6): e428-e432.
54
Liao Y, Meng Y, Liu H, et al. An automatic evaluation framework for multi-turn medical consultations capabilities of large language models [EB/OL].
55
Fast D, Adams LC, Busch F, et al. Autonomous medical evaluation for guideline adherence of large language models [J]. NPJ Digit Med, 2024, 7(1): 358-371.
56
Shi X, Xu J, Ding J, et al. LLM-mini-CEX: automatic evaluation of large language model for diagnosticconversation [EB/OL].
57
Busch F, Hoffmann L, Rueger C, et al. Current applications and challenges in large language models for patient care: a systematic review [J]. Commun Med, 2025, 5(1): 26-38.
58
Zhang B, Bornet A, Yazdani A, et al. A dataset for evaluating clinical research claims in large language models [J]. Sci Data, 2025, 12(1): 86.
59
US. Food and drug administration. FDA issues comprehensive draft guidance for developers of artificial intelligence-enabled medical devices [EB/OL].
60
Yang Y, Cui YU, Wang YT, et al. [Interpretation of the WHO's "ethics and governance of artificial intelligence for health: guidance on large multi-modal models" and its implications for China] [J]. Zhonghua Yu Fang Yi Xue Za Zhi, 2025, 59(6): 960-969.
61
Pashangpour S, Nejat G. The future of intelligent healthcare: a systematic analysis and discussion on the integration and impact of robots using large language models for healthcare [J]. Robotics, 2024, 13(8): 112.
62
Harvard University. Harvard designs AI sandbox that enables exploration, interaction without compromising security [EB/OL].
63
新加坡资讯通信媒体发展局. 生成式 AI 评估沙盒 [EB/OL]. (2023-10-31)[2026-06-20].
[1] 韩君, 徐兴祥. 智慧医疗在慢病管理中的研究进展[J/OL]. 中华肺部疾病杂志(电子版), 2020, 13(02): 268-270.
[2] 马木提江·木尔提扎, 汪永新, 阿西木江·阿西尔, 姜彦文, 秦虎. 多模态三维影像融合技术在颅内功能区病变手术中的应用[J/OL]. 中华神经创伤外科电子杂志, 2023, 09(05): 302-307.
[3] 李琦. 脑血管疾病半结构化专病病历及智慧医疗系统建立[J/OL]. 中华脑血管病杂志(电子版), 2021, 15(02): 108-111.
阅读次数
全文


摘要


AI


AI小编
你好!我是《中华医学电子期刊资源库》AI小编,有什么可以帮您的吗?