血浆ATR-FTIR光谱结合机器学习用于鼻咽癌分类。
Plasma ATR-FTIR spectroscopy combined with machine learning for nasopharyngeal carcinoma classification.
文献信息
| PMID | 42785009 |
|---|---|
| 原文 | 在 PubMed 查看原文 ↗ |
| 发表日期 | 2026 |
| 作者 | Rock Christian Tomas |
| 作者单位 | Department of Electrical Engineering, University of the Philippines Los Baños, Laguna, Philippines. Electronic address: rvtomas1@up.edu.ph. |
| 期刊 | Spectrochimica acta. Part A, Molecular and biomolecular spectroscopy |
| SCI 分区 | Q1 |
| IF | 4.4 |
| 研究类型 | AI/ML · 临床 |
| 所属专科 | 鼻咽癌 |
中文摘要
背景: 鼻咽癌(NPC)常在晚期才被诊断,因此需要微创方法来支持检测。本研究评估了血浆衰减全反射傅里叶变换红外(ATR-FTIR)光谱结合机器学习,以区分NPC与临床健康个体。
方法: 对51例经组织学确诊的NPC病例和51例年龄、性别匹配的临床健康对照者的血浆样本在4000-600 cm-1范围内进行分析。使用Mann-Whitney U检验评估21个光谱峰处的差异。使用全光谱、指纹区、选定峰和低秩光谱表示,通过重复交叉验证和探索性年龄、性别分层分析,评估了七种机器学习算法。
结果: 13个峰在中位吸光度上显示出显著的组间差异(p < 0.01)。尽管原始光谱分布存在大量重叠,但在完整队列中,当使用选定峰和低秩光谱表示时,模型性能有所提高:最佳情况AUC分别为全光谱0.6303 ± 0.0384、指纹区0.6330 ± 0.0333、选定峰0.7066 ± 0.0452、低秩表示0.7404 ± 0.0422。神经网络模型使用低秩特征取得了最佳总体性能,准确率(ACC)为0.7115 ± 0.0456。探索性亚组估计在年龄和性别定义的亚组间存在差异。然而,由于亚组样本量有限且性能估计存在变异性,对解释和普遍性需谨慎。
结论: 血浆ATR-FTIR光谱结合机器学习在NPC与临床健康对照者之间显示出中等程度的区分能力,使用低秩光谱特征和前馈神经网络获得最高性能。需要更大规模的独立队列进行验证。
英文摘要
BACKGROUND: Nasopharyngeal carcinoma (NPC) is often diagnosed at advanced stages, creating a need for minimally invasive approaches to support detection. This study evaluated plasma attenuated total reflectance Fourier transform infrared (ATR-FTIR) spectroscopy, combined with machine learning, to distinguish NPC from clinically healthy individuals.
METHODS: Plasma samples from 51 histologically confirmed NPC cases and 51 age- and sex-matched clinically healthy controls were analyzed across 4000-600 cm-1. Differences at 21 spectral peaks were assessed using the Mann-Whitney U test. Seven machine-learning algorithms were evaluated using the full spectrum, fingerprint region, selected peaks, and low-rank spectral representations with repeated cross-validation and exploratory age- and sex-stratified analyses.
RESULTS: Thirteen peaks showed significant between-group differences in median absorbance (p < 0.01). Despite substantial overlap in the original spectral distributions, model performance in the complete cohort increased when selected peaks and low-rank spectral representations were used: the best case AUCs were 0.6303 ± 0.0384 for the full spectrum, 0.6330 ± 0.0333 for the fingerprint region, 0.7066 ± 0.0452 for selected peaks, and 0.7404 ± 0.0422 for the low-rank representation. The neural network model achieved the best overall performance using low-rank features, with an accuracy (ACC) of 0.7115 ± 0.0456. Exploratory subgroup estimates varied across age- and sex-defined subsets. However, due to limited subgroup sizes and variability in the performance estimates, interpretation and generalizability was cautioned.
CONCLUSION: Plasma ATR-FTIR spectroscopy combined with machine learning showed moderate discrimination between NPC and clinically healthy controls, with the highest performance obtained using low-rank spectral features and a feedforward neural network. Larger independent cohorts are needed for validation.