基于Transformer的多模态融合框架用于接触式内窥镜喉部病变分类
A transformer-based multi-modal fusion framework for laryngeal lesion classification using contact endoscopy.
文献信息
| PMID | 42756136 |
|---|---|
| 原文 | 在 PubMed 查看原文 ↗ |
| 发表日期 | 2026 |
| 作者 | Yunjing Duan |
| 作者单位 | Department of Otolaryngology - Head and Neck Surgery, Shijiazhuang People's Hospital, Shijiazhuang, Hebei, China. |
| 期刊 | Frontiers in surgery |
| SCI 分区 | Q2 |
| IF | 2.1 |
| 研究类型 | AI/ML · 临床 |
| 所属专科 | 咽喉科 |
中文摘要
目的: 本研究介绍了一种新颖的基于Transformer的多模态框架,该框架整合了临床数据、影像组学特征和深度学习表示,用于从接触式内窥镜图像中自动分类喉部病变。
方法: 我们回顾性纳入了来自三个独立医疗中心(中心A、B和C)的300名喉部病变患者,获取了7,847张带有窄带成像的接触式内窥镜图像。临床变量(18个特征)、使用PyRadiomics提取的影像组学特征(156个特征)以及来自三种最先进架构(ConvNeXt V2、Swin Transformer V2和EfficientNet V2)的深度学习表示(共3,328个特征)通过一个具有多头交叉注意力机制的六层Transformer编码器系统整合。我们评估了三种特征选择策略(LASSO、互信息、ReliefF)和三种先进的分类架构(Vision Transformer、Attention-MLP、图神经网络)。该框架使用中心A和B的数据(n=230名患者)进行开发,并进行了严格的五折交叉验证,并在来自中心C的独立外部测试集(n=70名患者)上进行了验证。
结果: 最佳配置(LASSO特征选择与Vision Transformer分类器)在训练集上达到96.1±1.0%的准确率(AUC-ROC:0.992±0.007),在内部验证集上达到93.8±1.4%(AUC-ROC:0.981±0.012),在外部测试集上达到91.4±2.1%(AUC-ROC:0.967±0.019),外部队列的敏感性为90.0±2.8%,特异性为92.5±2.5%。多模态融合方法显著优于所有单模态方法。
方法: 仅临床特征(验证准确率71.3%)、仅影像组学特征(84.2%)以及最佳单个深度学习模型(87.6%),分别提高了22.5、9.6和6.2个百分点(所有p<0.001)。注意力机制分析显示,模型动态加权模态贡献,对于正确分类的恶性病例,将47.0%的注意力分配给深度特征,34.7%分配给影像组学特征,18.3%分配给临床特征。
结论: 本研究提出了首个通过基于Transformer的融合系统整合临床数据、影像组学和深度学习的喉部病变分类综合框架。
英文摘要
PURPOSE: This study introduces a novel multi-modal transformer-based framework that integrates clinical data, radiomic features, and deep learning representations for automated classification of laryngeal lesions from contact endoscopy images.
METHODS: We retrospectively enrolled 300 patients with laryngeal lesions from three independent medical centers (Centers A, B, and C), acquiring 7,847 contact endoscopy images with narrow band imaging. Clinical variables (18 features), radiomic features extracted using PyRadiomics (156 features), and deep learning representations from three state-of-the-art architectures, ConvNeXt V2, Swin Transformer V2, and EfficientNet V2 (3,328 combined features), were systematically integrated through a six-layer transformer encoder with multi-head cross-attention mechanisms. We evaluated three feature selection strategies (LASSO, mutual information, ReliefF) and three advanced classification architectures (Vision Transformer, Attention-MLP, Graph Neural Network). The framework was developed using data from Centers A and B (n = 230 patients) with rigorous five-fold cross-validation, and validated on an independent external test set from Center C (n = 70 patients).
RESULTS: The optimal configuration (LASSO feature selection with Vision Transformer classifier) achieved accuracy of 96.1 ± 1.0% (AUC-ROC: 0.992 ± 0.007) on training, 93.8 ± 1.4% (AUC-ROC: 0.981 ± 0.012) on internal validation, and 91.4 ± 2.1% (AUC-ROC: 0.967 ± 0.019) on external testing, with sensitivity of 90.0 ± 2.8% and specificity of 92.5 ± 2.5% on the external cohort. The multi-modal fusion approach significantly outperformed all single-modality.
METHODS: clinical features alone (71.3% validation accuracy), radiomic features alone (84.2%), and the best individual deep learning model (87.6%), with improvements of 22.5, 9.6, and 6.2 percentage points respectively (all p < 0.001. Attention mechanism analysis revealed that the model dynamically weighted modality contributions, allocating 47.0% attention to deep features, 34.7% to radiomic features, and 18.3% attention to clinical features for correctly classified malignant cases.
CONCLUSION: This study presents the first comprehensive framework for laryngeal lesion classification that systematically integrates clinical data, radiomics, and deep learning through transformer-based fusion.