基于人工智能的阻塞性睡眠呼吸暂停筛查工具在不同呼吸暂停低通气指数阈值下的性能:系统评价与荟萃分析
Performance of AI-Based Screening Tools for Obstructive Sleep Apnea Across Apnea-Hypopnea Index Thresholds: Systematic Review and Meta-Analysis.
文献信息
| PMID | 42727078 |
|---|---|
| 原文 | 在 PubMed 查看原文 ↗ |
| 发表日期 | 2026 |
| 作者 | Yujia Lv |
| 作者单位 | Department of Epidemiology and Health Statistics, School of Public Health, Tianjin Medical University, 22 Qixiangtai Road, Heping District, Tianjin, 300070, China, 86 022-83336619. |
| 期刊 | Journal of medical Internet research |
| SCI 分区 | Q1 |
| IF | 8.1 |
| 研究类型 | 综述 Meta · 临床 |
| 所属专科 | 鼻科 |
中文摘要
背景: 阻塞性睡眠呼吸暂停(OSA)患病率很高,但仍有大量患者未被诊断。多导睡眠监测(PSG)是参考标准,但其成本和有限的可及性限制了大规模病例识别。基于人工智能(AI)的筛查工具可能有助于风险分层和转诊优先级排序,但其在不同呼吸暂停低通气指数(AHI)阈值下的诊断准确性仍不确定。
目的: 本综述旨在系统评价基于AI的OSA筛查工具在AHI阈值≥5、≥15和≥30次/小时下的诊断准确性,重点关注使用非PSG衍生输入的模型。
方法: 检索了PubMed、Embase、Scopus和Web of Science,纳入2016年1月1日至2026年5月3日发表的研究。符合条件的研究包括:因疑似OSA接受评估或从人群队列中招募的成人;评估用于OSA筛查、风险预测或筛查导向的严重程度分类的AI模型;使用PSG作为参考标准;并报告了足以构建或重建2×2列联表的数据。采用双变量随机效应模型,按AHI阈值和输入来源分别汇总诊断准确性,并计算95%置信区间(CI)和预测区间(PI)。分别使用QUADAS-2(诊断准确性研究质量评估2)和GRADE(推荐分级的评估、制定与评价)评估偏倚风险和证据确定性。
结果: 共纳入60项研究,其中47项为荟萃分析提供了数据。在AHI阈值≥5、≥15和≥30次/小时下,合并敏感度分别为0.94(95% CI 0.92-0.96;95% PI 0.71-0.99)、0.87(95% CI 0.84-0.89;95% PI 0.66-0.96)和0.83(95% CI 0.79-0.87;95% PI 0.61-0.94);相应的特异度分别为0.77(95% CI 0.69-0.84;95% PI 0.30-0.96)、0.81(95% CI 0.75-0.85;95% PI 0.39-0.96)和0.91(95% CI 0.87-0.94;95% PI 0.55-0.99)。相应的汇总受试者工作特征曲线下面积分别为0.943、0.907和0.920。对于非PSG衍生工具,在3个阈值下的敏感度分别为0.92、0.85和0.81,特异度分别为0.70、0.74和0.85。对于PSG衍生模型,敏感度分别为0.96、0.90和0.85,特异度分别为0.82、0.88和0.96。探索性亚组分析提示,性能因所选研究和模型特征(包括地区、算法框架、数据来源和验证方法)而异。
结论: 基于AI的工具在临床相关的AHI阈值下对OSA显示出总体良好的筛查性能,尽管较宽的PI提示在未来可比人群和环境中性能可能变化。通过综合3个AHI阈值下的诊断准确性并区分非PSG衍生模型与PSG衍生模型,本综述扩展了既往宽泛或特定模态的综述,并为将模型性能与预期用途联系起来提供了临床可解释、路径特异性的依据。研究结果可能阐明非PSG衍生工具在前端筛查和转诊优先级排序中的潜在作用,以及PSG衍生模型在减少导联评估和睡眠实验室工作流程支持中的作用。鉴于存在显著异质性、外部验证有限以及证据确定性低或极低,在常规实施前需要进行前瞻性验证。
英文摘要
BACKGROUND: Obstructive sleep apnea (OSA) is highly prevalent but remains substantially underdiagnosed. Polysomnography (PSG) is the reference standard, but its cost and limited availability constrain large-scale case identification. AI-based screening tools may support risk stratification and referral prioritization, but their diagnostic accuracy across apnea-hypopnea index (AHI) thresholds remains uncertain.
OBJECTIVE: This review aimed to systematically evaluate the diagnostic accuracy of AI-based OSA screening tools at AHI thresholds of ≥5, ≥15, and ≥30 events/hour, with emphasis on models using non-PSG-derived inputs.
METHODS: PubMed, Embase, Scopus, and Web of Science were searched for studies published from January 1, 2016, to May 3, 2026. Eligible studies included adults evaluated for suspected OSA or recruited from population-based cohorts, assessed AI-based models intended or interpretable for OSA screening, risk prediction, or screening-oriented severity classification, used PSG as the reference standard, and reported sufficient data to construct or reconstruct 2×2 contingency tables. Diagnostic accuracy was synthesized separately by AHI threshold and input source using bivariate random-effects models, with 95% CIs and prediction intervals (PIs). Risk of bias and certainty of evidence were assessed using QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2) and GRADE (Grading of Recommendations Assessment, Development, and Evaluation), respectively.
RESULTS: A total of 60 studies were included, of which 47 contributed data to the meta-analysis. At AHI thresholds of ≥5, ≥15, and ≥30 events/hour, pooled sensitivities were 0.94 (95% CI 0.92-0.96; 95% PI 0.71-0.99), 0.87 (95% CI 0.84-0.89; 95% PI 0.66-0.96), and 0.83 (95% CI 0.79-0.87; 95% PI 0.61-0.94), respectively; the corresponding specificities were 0.77 (95% CI 0.69-0.84; 95% PI 0.30-0.96), 0.81 (95% CI 0.75-0.85; 95% PI 0.39-0.96), and 0.91 (95% CI 0.87-0.94; 95% PI 0.55-0.99), respectively. The corresponding areas under the summary receiver operating characteristic curves were 0.943, 0.907, and 0.920. For non-PSG-derived tools, sensitivities were 0.92, 0.85, and 0.81, and specificities were 0.70, 0.74, and 0.85 at the 3 thresholds, respectively. For PSG-derived models, sensitivities were 0.96, 0.90, and 0.85, and specificities were 0.82, 0.88, and 0.96, respectively. Exploratory subgroup analyses suggested performance variation across selected study and model characteristics, including region, algorithmic framework, data source, and validation method.
CONCLUSIONS: AI-based tools showed generally favorable screening performance for OSA across clinically relevant AHI thresholds, although wide PIs suggest variable performance across future comparable populations and settings. By synthesizing diagnostic accuracy across 3 AHI thresholds and distinguishing non-PSG-derived from PSG-derived models, this review extends previous broad or modality-specific reviews and offers a clinically interpretable, pathway-specific basis for linking model performance to intended use. The findings may clarify potential roles for non-PSG-derived tools in front-end screening and referral prioritization and for PSG-derived models in reduced-channel assessment and sleep-laboratory workflow support. Given substantial heterogeneity, limited external validation, and low or very low certainty of evidence, prospective validation is needed before routine implementation.