节点文献

基于多延迟四阶累积量倍频程谱线的腭裂语音咽擦音自动检测算法

Automatic Detection Algorithm of Pharyngeal Fricative in Cleft Palate Speech Based on Multi-delay Fourth-order Cumulant Octave Spectral Line

  • 推荐 CAJ下载
  • PDF下载
  • 不支持迅雷等下载工具,请取消加速工具后下载。

【作者】 何飞孟雨璇田维维王熙月何凌尹恒

【Author】 HE Fei;MENG Yu-xuan;TIAN Wei-wei;WANG Xi-yue;HE Ling;YIN Heng;School of Electrical Engineering and Information,Sichuan University;State Key Laboratory of Oral Diseases;

【通讯作者】 尹恒;

【机构】 四川大学电气信息学院口腔疾病研究国家重点实验室

【摘要】 为了实现对腭裂语音咽擦音及正常音节的自动分类检测,通过对腭裂咽擦患者发音特点的研究,提出了基于多延迟四阶累积量倍频程谱线(Fourth-order Cumulant One-third Octave Spectra Line,FTSL)的腭裂语音咽擦音自动检测算法。目前,咽擦音的研究多基于咽擦音的辅音时长及其在频域的能量分布等特征,实现了咽擦音及正常擦音自动检测的其他研究较少。文中实验基于腭裂语音咽擦音的发音特性,通过研究语音信号的多延迟四阶累计量,利用1/3倍频程算法提取特征谱线,实现了腭裂语音咽擦音与正常擦音的自动分类检测。实验提取了200个正常擦音辅音和194个腭裂语音咽擦音辅音的FTSL特征谱线,使用SVM(Support Vector Machine)分类器进行分类,并设计了FTSL谱线与其他传统语音特征的对比实验,进行了充分的分析讨论。实验结果表明,FTSL谱线对咽擦音的自动分类检测正确率高达92.7%,具有较优的性能,能为临床腭咽功能评估提供有效、客观、无创的辅助依据。

【Abstract】 In order to realize the automatic classification and detection of palate pharyngeal fricative and normal speech, an automatic pharyngeal fricative detection algorithm based on multi-delay fourth-order cumulant one-third octave spectral line(FTSL) was proposed by studying the pronunciation characteristics of cleft palate patients with pharyngeal fricative.Currently,most researches involved with the detection of pharyngeal fricatives are based on the length of consonants and the energy distribution of speech in frequency-domain.There exist few researches which have achieved automatic classification of pharyngeal fricatives and normal speech.This experiment is based on the pronunciation characteristics of pharyngeal fricative.Each frame’s multi-delay fourth-ordercumulant is computed,and then one-third octave is used to extract the FTSL.Automatic classification of pharyngeal fricative and normal speech is realized by FTSL.In this experiment,the FTSL of 200 normal consonants and 194 consonants of pharyngeal fricative are extracted,and the SVM classifier is used to classify.Besides,comparative experiments were conducted on FTSL feature and traditional acoustic features,and the results were fully analyzed and discussed in this paper.The experimental results show that the proposed FTSL has an accurate rate of 92.7% for the automatic classification of pharyngeal speeches,and it has excellent performance and can provide an effective,objective and non-invasive auxiliary basis for clinical pharyngeal state assessment.

【基金】 国家自然基金青年科学基金(61503264)~~
  • 【文献出处】 计算机科学 ,Computer Science , 编辑部邮箱 ,2020年01期
  • 【分类号】R767.92;TN912.3
  • 【被引频次】4
  • 【下载频次】104
节点文献中: 

本文链接的文献网络图示:

本文的引文网络