药物诱导睡眠内镜的评分者间信度:经验驱动的解读、决策以及人工智能的新兴作用。
Inter-rater reliability in drug-induced sleep endoscopy: experience-driven interpretation, decision-making, and the emerging role of artificial intelligence.
文献信息
| PMID | 42766212 |
|---|---|
| 原文 | 在 PubMed 查看原文 ↗ |
| 发表日期 | 2026 |
| 作者 | Federico Leone |
| 作者单位 | Department of Otorhinolaryngology, Sleep Surgery Center, Sleep Disorders Center, Istituto Auxologico Italiano IRCCS, Milan, Italy. doc.federicoleone@gmail.com. |
| 期刊 | Sleep & breathing = Schlaf & Atmung |
| SCI 分区 | Q3 |
| IF | 2.8 |
| 研究类型 | AI/ML · 临床 |
| 所属专科 | 咽喉科 |
中文摘要
目的: 评估使用VOTE分类对DISE解读的评分者间信度,评估治疗决策的变异性,并探索人工智能(AI)在DISE评估中的潜在作用。
方法: 20份DISE记录由12名评分者(7名初级和5名高级)以及一个未经任务特定训练的通用AI模型进行回顾性独立评估。使用百分比一致性、Fleiss' kappa和加权kappa评估人类评分者之间在阻塞分级和模式分类方面的一致性。分析了与参考标准(专家一致共识)的一致性。使用精确匹配率和Jaccard相似度评估治疗建议。
结果: 在基线时,人类评分者之间在阻塞分级方面的一致性范围为43.6%至68.0%(加权κ,0.255-0.425),在塌陷模式方面为60.8%至79.3%(Fleiss' κ,0.287-0.370)。高级评分者在所有解剖水平的塌陷模式以及四个水平中的三个的阻塞分级方面显示出比初级评分者更高的基线一致性。与专家共识的一致性在基线分级中,高级评分者为70.0%至90.0%,AI为35.0%至85.0%;在基线模式分类中,分别为75.0%至95.0%和50.0%至80.0%。对于治疗建议,高级评分者的精确一致性为70.0%,初级评分者为55.0%,AI为50.0%;平均Jaccard相似度分别为0.867、0.833和0.742。
结论: DISE解读仍然复杂且依赖于操作者。模式分类似乎比分级更具可重复性。虽然非任务特定的AI模型捕捉到了临床相关模式,但在复杂决策中未能达到专家表现。AI可能作为一种支持工具,特别是对于经验较少的临床医生。
英文摘要
PURPOSE: To evaluate the inter-rater reliability of DISE interpretation using the VOTE classification, assess variability in therapeutic decision-making, and explore the potential role of artificial intelligence (AI) in DISE evaluation.
METHODS: Twenty DISE recordings were retrospectively and independently evaluated by 12 raters (7 junior and 5 senior) and by a general-purpose AI model without task-specific training. Inter-rater agreement among human raters for obstruction grading and pattern classification was assessed using percentage agreement, Fleiss' kappa, and weighted kappa. Agreement with a reference standard (unanimous expert consensus) was analyzed. Treatment proposals were evaluated using exact match ratio and Jaccard similarity.
RESULTS: At baseline, agreement among human raters ranged from 43.6% to 68.0% for obstruction grading (weighted κ, 0.255-0.425) and from 60.8% to 79.3% for collapse pattern (Fleiss' κ, 0.287-0.370). Senior raters showed higher baseline agreement than junior raters across all anatomical levels for collapse pattern and at three of four levels for obstruction grading. Agreement with the expert consensus ranged from 70.0% to 90.0% for senior raters and from 35.0% to 85.0% for AI in baseline grading, and from 75.0% to 95.0% and 50.0% to 80.0%, respectively, in baseline pattern classification. For therapeutic recommendations, exact agreement was 70.0% for senior raters, 55.0% for junior raters, and 50.0% for AI; mean Jaccard similarity was 0.867, 0.833, and 0.742, respectively.
CONCLUSIONS: DISE interpretation remains complex and operator-dependent. Pattern classification appears more reproducible than grading. While a non-task-specific AI model captured clinically relevant patterns, it did not match expert performance in complex decision-making. AI may serve as a supportive tool, particularly for less experienced clinicians.