FUDAN VOICE的验证:与计算机语音实验室(CSL)的信号和声学特征比较
Validation of FUDAN VOICE: Signal and Acoustic Feature Comparison with Computerized Speech Lab (CSL).
文献信息
| PMID | 42760223 |
|---|---|
| 原文 | 在 PubMed 查看原文 ↗ |
| 发表日期 | 2026 |
| 作者 | Siyan Xu |
| 作者单位 | ENT Institute and Department of , Eye & ENT Hospital, Fudan University, 83 Fenyang Road, Shanghai 200031, China. Electronic address: siyanx1120@163.com. |
| 期刊 | Journal of voice : official journal of the Voice Foundation |
| SCI 分区 | Q1 |
| IF | 2.4 |
| 研究类型 | 临床研究 · 临床 |
| 所属专科 | 咽喉科 |
中文摘要
目的: 本研究旨在验证自定义微信小程序FUDAN VOICE(在我们早期研究中命名为Voice Acquisition)采集的音频与计算机语音实验室(CSL)录音在信号和声学特征水平上的一致性,并描述两种录音路径之间的总体信号偏差特征,涵盖硬件、平台级处理和环境噪声的贡献。
方法: 采用同步录音设计,在隔音室和安静办公环境中同时采集CSL和FUDAN VOICE的平行音频。使用均方根误差(RMSE)、皮尔逊相关系数(Pearson's rho)和平均相干性评估信号水平的一致性。提取基频(F0)、抖动、闪烁、噪声谐波比(NHR)、倒谱峰值突出度(CPP)以及倒谱/频谱发声困难指数(CSID)等声学特征,并进行相关分析和组内相关系数(ICC)比较。
结果: 信号域分析显示FUDAN VOICE与CSL之间具有良好的一致性,RMSE较低,且在5 kHz以下保留了频谱结构,但观察到高频衰减(> 5 kHz)。特征级验证表明,在不同环境中,F0均值、CPP和CSID具有极好的一致性(r > 0.82,ICC > 0.82),而F0极端值——尤其是F0最小值——在非隔音环境中表现出显著退化。抖动和NHR保持稳健,而闪烁则表现出环境敏感性。
结论: FUDAN VOICE在隔音室和办公环境中均能实现可靠的远程声学采集,适用于F0均值、CPP和CSID,但高频衰减(> 5 kHz)和噪声环境下的F0退化值得注意。本研究建立了首个基于超级应用的语音采集验证框架,支持微信小程序在移动健康中的标准化。
英文摘要
OBJECTIVE: This study aimed to validate the consistency between audio acquired by a custom WeChat mini-program, FUDAN VOICE (named Voice Acquisition in our earlier study), and Computerized Speech Lab (CSL) recordings at both signal and acoustic feature levels, and to characterize aggregate signal deviations between the two recording pathways, encompassing contributions from hardware, platform-level processing, and environmental noise.
METHODS: A simultaneous recording design was employed, capturing parallel audio from CSL and FUDAN VOICE across a soundproof room and a quiet office environment. Signal-level agreement was assessed using the root mean square error (RMSE), Pearson correlation coefficient (Pearson's rho), and mean coherence. Acoustic features such as fundamental frequency (F0), jitter, shimmer, noise-to-harmonic ratio (NHR), cepstral peak prominence (CPP), and cepstral/spectral index of dysphonia (CSID) were extracted and compared for correlation analysis and intraclass correlation coefficients (ICC).
RESULTS: Signal-domain analysis revealed favorable agreement between FUDAN VOICE and CSL, with low RMSE and preserved spectral structure below 5 kHz, although high-frequency attenuation (> 5 kHz) was observed. Feature-level verification demonstrated excellent concordance for F0 mean, CPP, and CSID (r > 0.82, ICC > 0.82) across environments, whereas F0 extreme values-particularly F0 min-showed marked degradation in non-soundproof settings. Jitter and NHR remained robust, while shimmer exhibited environmental sensitivity.
CONCLUSIONS: FUDAN VOICE achieves reliable remote acoustic acquisition for F0 mean, CPP, and CSID across soundproof and office environments, though high-frequency attenuation (> 5 kHz) and noisy-setting F0 degradation warrant caution. This study establishes the first validation framework for super-app-based voice acquisition, supporting the standardization of WeChat mini-programs in mobile health.