节点文献

大语言模型在神经眼科中应用的多中心评价

A multicenter evaluation study of the use of large language models in neuro-ophthalmology

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 王子荀; 张晓玲; 贾洪强; 魏瑞华; 王宇航; 范珂; 祁艳华; 谢学说; 魏世辉; 李志清;

【Author】 WANG Zixun;ZHANG Xiaoling;JIA Hongqiang;WEI Ruihua;WANG Yuhang;FAN Ke;QI Yanhua;XIE Xueshuo;WEI Shihui;LI Zhiqing;Tianjin Key Laboratory of Retinal Functions and Diseases,Tianjin Branch of National Clinical Research Center for Ocular Disease,Eye Institute and School of Optometry,Tianjin Medical University Eye Hospital;Handan City Eye Hospital(the Third Hospital of Handan);Cangzhou Eye Hospital;Department of Ophthalmology,Third Medical Center,PLA General Hospital;Henan Provincial People’s Hospital,Henan Provincial Eye Hospital;Lixiang Eye Hospital of Soochow University;Haihe Lab of ITAI;

【通讯作者】 李志清;

【机构】 天津医科大学眼科医院、眼视光学院、眼科研究所,国家眼耳鼻喉疾病临床医学研究中心天津市分中心,天津市视网膜功能与疾病重点实验室; 河北省邯郸市眼科医院,邯郸市第三医院; 沧州市眼科医院; 解放军总医院第三医学中心眼科医学部; 河南省人民医院,河南省立眼科医院; 苏州大学理想眼科医院; 先进计算与关键软件(信创)海河实验室;

【摘要】 目的 评价人工智能(AI)大语言模型(LLM)生成的与神经眼科相关典型临床问题的答案,并利用客观评价及专家评估的方式多维度探究神经眼科相关问题在LLM上的表现。方法 多中心、随机、横断面试验研究。从神经眼科疾病定义、病因、临床表现及体征检查和治疗及预后4个角度选取30个神经眼科领域相关典型问题,分别使用Deepseek、文心一言4.0、豆包及Kimi 1.5四种国内开源LLM输出答案文本,采取客观评估法定量分析;同时采取专家评估法,由三位眼科专家分别对120个答案文本进行量化评分。根据问题回答的完整性、准确性和专业性及相关性和实用性分别制定3级、5级及4级李克特量表。选取其中表现最佳的LLM,观察在4类问题中该LLM是否存在表现差异,由另外三位专家评估各LLM是否可以替代真实世界医患沟通。结果 在客观的汉语文本阅读难度分析中,4种LLM的总字数组间差异存在统计学意义(均为P<0.001)。在4种LLM中,Kimi 1.5表现最出色,其完整性最高分(3分)、准确性和专业性最高分(5分)、相关性和实用性最高分(4分)的频率分别为61%、29%、41%。Kimi 1.5在神经眼科疾病定义、病因、临床表现及体征、治疗及预后四个方面的问题中表现较为一致,组间差异均无统计学意义(均为P>0.05)。结论 中文LLM在神经眼科临床应用中具有较大的潜力,Kimi 1.5在完整性、准确性和专业性及相关性和实用性方面表现较其他LLM出色,但仍无法代替真实世界医患沟通,未来需要探索AI+医师的新型诊疗模式。

【Abstract】 Objective To evaluate answers to typical clinical questions related to neuro-ophthalmology generated by Artificial Intelligence(AI) Large Language Models(LLM) and to explore the performance of neuro-ophthalmology-related questions on LLM in a multidimensional manner using objective and expert assessment. Methods Multicenter, randomized, cross-sectional pilot study. Thirty typical questions related to neuro-ophthalmology were selected based on four perspectives: definition, etiology, clinical manifestations and signs, and treatment and prognosis, and were analyzed quantitatively using Deepseek, Wenxin Yiyin 4.0, Doubao, and Kimi 1.5, which are four open-source LLMs in China, and quantitatively analyzed with objective assessment; and quantitatively rated by three ophthalmologists using expert assessment for 120 answer texts. Three ophthalmology experts quantitatively scored the 120 answer texts. Three ophthalmologists quantitatively scored the 120 answer texts. Level 3, 5, and 4 Likert scales were developed according to the completeness, accuracy, professionalism, relevance, and criticality of the question texts, respectively. The best-performing LLM was selected, and its performance was observed across the four types of questions. Additionally, three other experts assessed whether the best-performing one could be evaluated as a substitute for real-world doctor-patient communication.Results In the objective Chinese text reading difficulty analysis, the differences in total word count among the four LLMs were statistically significant(all P<0.001). Of the four LLMs, Kimi 1.5 performed the best, with frequencies of 61%, 29%, and 41% for the highest scores in completeness(3), accuracy and professionalism(5), and relevance and usefulness(4), respectively. Kimi 1.5 performed more consistently on the questions on the four areas of neuro-ophthalmologic disorders: definition, etiology, clinical manifestations and signs, treatment, and prognosis, with no between-group differences(P>0.05). Conclusion Chinese language LLMs have great potential in the clinical application of neuro-ophthalmology. Kimi 1.5 outperforms other LLMs in terms of completeness, accuracy, professionalism, relevance, and usefulness, but it still cannot replace real-world doctor-patient communication. There is a need to explore new diagnostic and therapeutic model of AI+physician in the future.

【基金】 天津市医学重点学科建设资助项目(编号:TJYXZDXK-3-004A-2);天津市视网膜功能与疾病重点实验室自主与开放课题(编号:2023tjswmm004);天津医科大学眼科医院高水平创新型人才培养基金(编号:YDYYRCXM-B2023-02)
  • 【文献出处】 眼科新进展 ,Recent Advances in Ophthalmology , 编辑部邮箱 ,2025年10期
  • 【分类号】R77;TP18
  • 【下载频次】41
节点文献中: 

本文链接的文献网络图示:

本文的引文网络