耳鼻喉科Pubmed文献追踪每日 14:00 同步
← 返回全部文献
Article record

教室中听障儿童的语音分离

Speech separation for hearing-impaired children in the classroom.

AI/ML耳科IF 2.6Q1

文献信息

中文摘要

教室环境对听障儿童构成重大挑战,背景噪声、同时说话者和混响会降低语音感知。大多数基于深度学习的语音分离算法是针对简化条件下的成人语音开发的,忽视了儿童语音更高的频谱相似性以及教室的声学复杂性。我们使用MIMO-TasNet来填补这一空白,这是一种紧凑、低延迟、多通道架构,适用于双侧助听器或人工耳蜗中的实时处理。我们模拟了自然教室条件,包括移动说话者、儿童-儿童和儿童-成人配对,以及不同的噪声和距离设置。我们比较了三种训练策略:仅成人语音、教室特定数据,以及用有限的教室数据微调成人训练模型。结果表明,双耳空间线索使成人训练模型在干净的教室条件下表现良好,即使在重叠的儿童说话者上也是如此。教室特定训练提高了分离质量,而仅用一半数据微调就达到了更优性能,证实了数据高效适应的益处。使用扩散嘈杂噪声训练增强了跨条件的鲁棒性,模型在泛化到未见过的说话者-听者距离时保持了空间意识。将空间感知架构与针对性适应策略相结合,是增强儿童计算语音分离的重要一步,有助于指导针对教室噪声中语音挑战的实际设备辅助技术的部署。

英文摘要

Classroom environments pose significant challenges for hearing-impaired children, where background noise, simultaneous talkers, and reverberation degrade speech perception. Most deep learning-based speech separation algorithms are developed for adult voices in simplified conditions, neglecting both the higher spectral similarity of children's voices and the acoustic complexity of classrooms. We address this gap using MIMO-TasNet, a compact, low-latency, multi-channel architecture suited for real-time processing in bilateral hearing aids or cochlear implants. We simulated naturalistic classroom conditions with moving talkers, both child-child and child-adult pairs across varying noise and distance settings. We compared three training strategies: adult speech only, classroom-specific data, and finetuning adult-trained models with limited classroom data. Results show that binaural spatial cues enable adult-trained models to perform well in clean classroom conditions, even on overlapping child talkers. Classroom-specific training improved separation quality, while finetuning with only half the data achieved superior performance, confirming data-efficient adaptation benefits. Training with diffuse babble noise enhanced robustness across conditions, and models preserved spatial awareness while generalizing to unseen talker-listener distances. Combining spatially aware architectures with targeted adaptation strategies is an important step toward enhancing computational speech separation for children, which can help guide the deployment of practical on-device assistive technologies for classroom speech-in-noise challenges.