使用项目级SNOT-22反应预测90天内鼻科手术:一项纳入35,170例患者的多中心机器学习研究
Prediction of Rhinologic Surgery Within 90 Days Using Item-Level SNOT-22 Responses: A Multisite Machine Learning Study of 35,170 Patients.
文献信息
| PMID | 42722008 |
|---|---|
| 原文 | 在 PubMed 查看原文 ↗ |
| 发表日期 | 2026 |
| 作者 | Michael Sramek |
| 作者单位 | Department of Otolaryngology-Head and Neck Surgery, Mayo Clinic Arizona, Phoenix, Arizona, USA. |
| 期刊 | International forum of allergy & rhinology |
| SCI 分区 | Q1 |
| IF | 4.9 |
| 研究类型 | AI/ML · 临床 |
| 所属专科 | 鼻科 |
中文摘要
背景: 在容量受限的医疗系统中,优化对具有成本效益的医疗服务的可及性是一项当代的迫切任务。我们评估了使用患者报告数据的机器学习(ML)模型是否可以在无需CT成像的情况下,优化鼻科门诊需要手术治疗患者的可及性。结局定义为初次评估后90天内的任何鼻科手术。
方法: 使用来自一个整合医疗系统内五个不同地点、2018年至2025年间就诊患者的去标识化电子数据集来训练模型。研究了人口统计学数据和22项鼻窦结局测试(SNOT-22)反应。模型采用分层5折交叉验证进行训练,其中80%为开发队列,并在20%的留出验证队列中进行验证。超参数使用Optuna进行优化。主要结局为初次SNOT-22后90天内接受任何鼻科手术干预的表现。
结果: 评估了35,170例患者的数据。在模型中,项目级反应优于使用SNOT-22总分。在各模型中,XGBoost表现出最佳区分度,AUC为0.70(95% CI,0.69-0.71),优于逻辑回归(0.66)、随机森林(0.63)和TabNet(0.66)(均p < 0.001)。在最佳阈值下,XGBoost达到66%的敏感性、64%的特异性、28%的PPV和90%的NPV。手术的主要预测因素为年龄、鼻塞、面部疼痛或压迫感以及嗅觉或味觉减退。在留出队列上的验证保持稳定(AUC 0.70),尽管各地点手术率从11.5%到27.4%不等,仍具有较强区分度。
结论: 人口统计学数据和项目级SNOT-22反应成功开发了一个ML模型,该模型对接下来90天内是否进行手术表现出中等区分度和高NPV(90%)。与人工监督相平衡的ML模型可能加速分诊并优化鼻科门诊的手术产出。
英文摘要
BACKGROUND: Optimizing access to cost-effective care in capacity-constrained health systems is a contemporary imperative. We evaluated whether machine learning (ML) models using patient-reported data could optimize access for patients requiring surgical care in rhinology clinics, without need for CT imaging. The outcome was defined as any rhinologic surgery within 90 days of initial evaluation.
METHODS: A de-identified electronic dataset from patients seen at five distinct sites within an integrated healthcare system between 2018 and 2025 was used to train models. Demographic data and 22-item Sinonasal Outcome Test (SNOT-22) responses were studied. Models were trained using stratified 5-fold cross-validation with an 80% development cohort and validated in a 20% held-out validation cohort. Hyperparameters were optimized with Optuna. The primary outcome was performance of any rhinologic surgical intervention within 90 days of initial SNOT-22.
RESULTS: Data from 35,170 patients were evaluated. Item-level responses outperformed use of total SNOT-22 score in the models. Among models, XGBoost demonstrated best discrimination with AUC of 0.70 (95% CI, 0.69-0.71), outperforming logistic regression (0.66), random forest (0.63), and TabNet (0.66) (all p < 0.001). At optimal threshold, XGBoost achieved 66% sensitivity, 64% specificity, 28% PPV, and 90% NPV. Top predictors for surgery were age, nasal blockage, facial pain or pressure, and decreased sense of smell or taste. Validation on the held-out cohort remained stable (AUC 0.70), with strong discrimination across sites despite surgical rates ranging from 11.5% to 27.4%.
CONCLUSIONS: Demographics and item-level SNOT-22 responses were successful in developing an ML model that demonstrated moderate discrimination and high NPV (90%) for performance of surgery within next 90 days. ML models balanced with human oversight may accelerate triage and optimize surgical yield for rhinology clinics.