耳鼻喉科Pubmed文献追踪每日 14:00 同步
← 返回全部文献
Article record

利用GPT-5进行梅尼埃病内淋巴积水MR分级的人机协作

Human-AI collaboration using GPT-5 for MR grading of endolymphatic hydrops in Ménière's disease.

AI/ML耳科IF 5.7Q1

文献信息

中文摘要

目的: 评估GPT-5在梅尼埃病(MD)延迟钆增强3D-FLAIR MRI上对耳蜗和前庭内淋巴积水(EH)进行分级的能力,并评估提示策略和人机协作。
材料与方法: 这项回顾性研究纳入436例MD患者(872只耳)。GPT-5在多种提示策略下对耳蜗和前庭EH进行分级,并在独立工作流程和GPT-5辅助工作流程中与初级和高级神经放射科医生进行比较。评估了准确率(ACC)、AUC、F1评分、一致性和专家评定的临床实用性。
结果: 在无结构化图像描述的基本提示策略下,诊断性能仍然较差:尽管纳入了临床病史或少样本示例,耳蜗积水分级ACC仅达到25%,前庭积水分级ACC仅达到39%。相比之下,在整合了结构化图像描述、临床背景和代表性示例的增强提示下,GPT-5表现出显著改善的性能,耳蜗分级准确率达到87%,前庭分级准确率达到80%,一致性改善且专家评分良好。在人机交互实验中,Human-first模式改善了初级医生的表现,使耳蜗分级ACC从36%提高到59%,前庭分级ACC从42%提高到53%。
结论: 当主要依赖图像输入时,GPT-5独立对MRI上的EH进行分级的能力有限。只有在纳入结构化图像描述、临床背景和代表性示例后,才达到可接受的诊断性能。在这些增强提示条件下,GPT-5显示出作为人机协作判读辅助工具的潜力,尤其适用于经验较少的医生,而非作为完全自主的诊断系统。
要点: 问题:在3D-FLAIR MRI上准确分级内淋巴积水对于管理梅尼埃病至关重要,但当前判读缺乏标准化和效率。发现:在没有图像描述的情况下,GPT-5的EH MRI分级准确性较差,而优化提示和人机协作显著提高了性能,使其接近专家判读。临床相关性:GPT-5对内淋巴积水进行可接受的MRI分级需要结构化图像描述和情境化提示。这支持其作为人机协作辅助工具而非自主系统的作用,可能减少对有限医疗资源的依赖。

英文摘要

OBJECTIVE: To evaluate GPT-5 for grading cochlear and vestibular endolymphatic hydrops (EH) on delayed gadolinium-enhanced 3D-FLAIR MRI in Ménière's disease (MD) and to assess prompting strategies and human-AI collaboration.
MATERIALS AND METHODS: This retrospective study included 436 patients with MD (872 ears). GPT-5 graded cochlear and vestibular EH under multiple prompting strategies and was compared with junior and senior neuroradiologists in independent and GPT-5-assisted workflows. Accuracy (ACC), AUC, F1 score, agreement, and expert-rated clinical usefulness were assessed.
RESULTS: Under basic prompting strategies without structured image descriptions, diagnostic performance remained poor: cochlear hydrops grading ACC reached only 25%, and vestibular hydrops grading ACC reached only 39%, despite the inclusion of clinical history or few-shot examples. In contrast, with enhanced prompting that integrated structured image descriptions, clinical context, and representative examples, GPT-5 showed substantially improved performance, achieving accuracies of 87% for cochlear grading and 80% for vestibular grading, with improved agreement and favorable expert ratings. In human-AI interaction experiments, the Human-first mode improved junior physicians' performance, increasing cochlear grading ACC from 36% to 59% and vestibular grading ACC from 42% to 53%.
CONCLUSION: GPT-5 demonstrated limited capability to independently grade EH on MRI when relying primarily on image input. Acceptable diagnostic performance was achieved only after incorporating structured image descriptions, clinical context and representative examples. Under these enhanced prompting conditions, GPT-5 showed potential as an assistive tool for human-AI collaborative interpretation, particularly for less experienced physicians, rather than as a fully autonomous diagnostic system.
KEY POINTS: Question Accurate grading of endolymphatic hydrops on 3D-FLAIR MRI is critical for managing Ménière's disease, but current interpretation lacks standardization and efficiency. Findings GPT-5 showed poor EH MRI grading accuracy without image descriptions, while optimized prompting and human-AI collaboration markedly improved performance toward expert interpretation. Clinical relevance Acceptable MRI grading of endolymphatic hydrops by GPT-5 requires structured image descriptions and contextual prompting. This supports its role as a human-AI collaborative assistive tool, not an autonomous system, potentially reducing reliance on limited medical resources.