使用卷积神经网络对有吞咽困难和无吞咽困难个体吞咽声学信号频谱图进行分类的方法
Method for Classifying Spectrograms of Acoustic Signals from Swallowing of Individuals with and without Dysphagia Using a Convolutional Neural Network.
文献信息
| PMID | 42722902 |
|---|---|
| 原文 | 在 PubMed 查看原文 ↗ |
| 发表日期 | 2026 |
| 作者 | Gabriele Pessoa da Silva |
| 作者单位 | Graduate Program in Information Technology and Health Management, Federal University of Health Sciences of Porto Alegre, Porto Alegre, Rio Grande do Sul, Brazil. gabriele.pessoa.silva@gmail.com. |
| 期刊 | Dysphagia |
| SCI 分区 | Q1 |
| IF | 3.1 |
| 研究类型 | AI/ML · 临床 |
| 所属专科 | 咽喉科 |
中文摘要
吞咽困难非常普遍,并可能导致重要的临床并发症。尽管传统诊断方法有效,但其存在局限性,包括高成本和患者不适。颈部听诊是一种无创监测吞咽声音的技术,可能支持吞咽困难评估。介绍一种使用机器学习方法分析和分类通过颈部听诊获得的吞咽困难个体和无吞咽困难个体声学信号的方法。横断面研究,纳入18岁及以上个体的吞咽评估。在临床吞咽困难评估的同时,使用与3 M-LITTMANN听诊器耦合的Eko Core放大器记录吞咽声音。听诊器置于参与者环状软骨正下方气管侧缘,并指示参与者在记录期间不要说话或发出不必要的声音,以尽量减少背景噪声。使用数字听诊器捕获吞咽困难个体和无吞咽困难个体的声学信号并存储以供分析。共选择178个吞咽信号,分割并增强为1,888个孤立事件,使用短时傅里叶变换将其转换为频谱图。使用的学习模型是卷积神经网络结合交叉验证系统。该模型达到77.7%的准确率和76.6%的灵敏度,但精确率相对较低(64.1%)。ROC曲线下面积为0.823。本研究通过表明神经网络支持的声学分析可能补充传统方法,推进了吞咽困难评估的自动化。未来的研究应侧重于提高模型的稳健性和临床适用性,以支持更易获得且侵入性更小的诊断方法。
英文摘要
Dysphagia is highly prevalent and may lead to important clinical complications. Although traditional diagnostic methods are effective, they have limitations, including high cost and patient discomfort. Cervical auscultation is a non-invasive technique for monitoring swallowing sounds and may support dysphagia assessment. To present a method for analyzing and classifying acoustic signals obtained through cervical auscultation of individuals with and without dysphagia using a machine learning approach. Cross-sectional study including swallowing assessments from individuals aged 18 years or older. Simultaneously with the clinical assessment of dysphagia, swallowing sounds were recorded using an Eko Core amplifier coupled to a 3 M-LITTMANN stethoscope. The stethoscope was positioned over the lateral edge of the trachea just below the participant's cricoid cartilage, and the participant was instructed not to speak or make unnecessary sounds during the recording to minimize background noise. Acoustic signals from individuals with and without dysphagia were captured with a digital stethoscope and stored for analysis. A total of 178 swallowing signals were selected, segmented, and augmented to 1,888 isolated events, which were converted into spectrograms using the Short-Time Fourier Transform. The learning model used was a convolutional neural network associated with a cross-validation system. The model achieved an accuracy of 77.7% and a sensitivity of 76.6%, but relatively low precision (64.1%). The area under the ROC curve was 0.823. This study advances the automation of dysphagia assessment by showing that acoustic analysis supported by neural networks may complement traditional methods. Future investigations should focus on improving model robustness and clinical applicability to support a more accessible and less invasive diagnostic approach.