耳鼻喉科Pubmed文献追踪每日 14:00 同步
← 返回全部文献
Article record

基于声带运动学的功能性喉部行为实时分类

Real-Time Classification of Functional Laryngeal Behaviors From Vocal Fold Kinematics.

AI/ML咽喉科IF 2.3Q2

文献信息

中文摘要

目的: 开发一种实时预测模型,用于在临床检查过程中从基于视频喉镜的追踪数据中检测喉部行为。这将通过为临床医生提供即时反馈来提高数据收集质量,并为未来实现高效大规模分析的自动化流程奠定基础。
方法: 我们训练了一个有状态的残差门控循环单元(GRU)模型,该模型分析来自我们已发表的关键点检测模型所导出的39个喉部关键点的追踪数据,以预测患者的喉部任务状态。这些状态包括“发声”、“持续发声”、“吞咽”、“空闲”、“咳嗽”、“嗅闻”和“视野外”。模型开发使用了来自72个喉镜视频的916个状态片段,包含222,065帧视频。所得模型在一个独立数据集上进行了评估,该数据集包含8个视频、123个片段和49,770帧。性能通过分类指标和时间交并比(mIoU)进行评估。
结果: 在独立测试数据集上评估时,该模型达到了92%的平均准确率和0.82的mIoU,表明与人工标注一致,并能准确地在时间上识别喉部任务状态。在类别层面,模型在大多数状态类别上表现一致,测试集上各类别的F1分数从发声的83%到嗅闻的97%不等。验证集上最低的F1出现在咳嗽(82% [70%-92%])。
结论: 在临床检查中,基于视频喉镜姿态追踪的喉部状态实时分类是可行的。通过对吞咽和发声等状态的可靠自动检测,这项工作为进一步自动化分析、解释和记录喉部病理奠定了基础。

英文摘要

OBJECTIVES: Develop a real-time prediction model to detect laryngeal behaviors from videolaryngoscopy-based tracking data during clinical examinations. This will improve data collection quality by providing immediate feedback to clinicians and enable future automation pipelines for efficient large-scale analysis.
METHODS: We trained a stateful residual Gated Recurrent Unit (GRU) model that analyzed the tracking data of 39 laryngeal keypoints, derived from our published keypoint detection model, to predict the patient's laryngeal task state. These included "phonation," "sustained phonation," "swallowing," "idle," "coughing," "sniffing," and "out of view." Model development used 916 state segments comprising 222,065 video frames from 72 laryngoscopy videos. The resulting model was evaluated on an independent dataset comprising 8 videos, 123 segments, and 49,770 frames. Performance was assessed using classification metrics and temporal intersection over union (mIoU).
RESULTS: The model achieved a mean accuracy score of 92% and an mIoU of 0.82 when evaluated on the independent test dataset, indicating agreement with manual annotations and accurate temporal identification of laryngeal task states. At the class level, the model performed consistently across most state categories, with per-class F1-scores on the test set ranging from 83% for phonation to 97% for sniffing. The lowest validation F1 was observed for cough (82% [70%-92%]).
CONCLUSION: Real-time classification of laryngeal states from videolaryngoscopy-based pose tracking during clinical examinations is feasible. With reliable automated detection of states such as swallowing and phonation, this work establishes a foundation for further automated analysis, interpretation, and documentation of laryngeal pathology.