节点文献
基于NMF和FCRF的单通道语音分离算法
Single-channel Speech Separation with Non-negative Matrix Factorization and Factorial Conditional Random Field
【作者】 李煦; 屠明; 吴超; 国雁萌; 纳跃跃; 付强; 颜永红;
【Author】 Li Xu;Tu Ming;Wu Chao;Guo Yan Meng;Na YueYue;Fu Qiang;Yan Yonghong;Key Laboratory of Speech Acoustics and Content Understanding Institute of Acoustics,Chinese Academy of Sciences;Signal Analysis Representation and Perception Laboratory, Arizona State University;
【机构】 中国科学院声学研究所语言声学与内容理解重点实验室; Signal Analysis Representation and Perception Laboratory,Arizona State University;
【摘要】 近年来,非负矩阵分解(Non-negative matrix factorization,NMF)被广泛应用于单通道语音分离问题。然而,标准的NMF算法假设语音的相邻帧之间是相互独立的,不能表征语音信号的时间连续性信息。为此,本文提出了一种新的语音分离算法,首先将NMF和k均值聚类结合对纯净语音的频谱结构以及时间连续性进行建模,然后利用得到的模型训练因子条件随机场(factorial conditional random field,FCRF),进而对混合语音信号进行分离。结果表明本文提出的算法相比于没有考虑语音时间连续特性的基于NMF的算法,如Active-Set Newton Algorithm(ASNA),在客观指标上有明显提高。
【Abstract】 Recently, Non-negative matrix factorization(NMF) has been extensively used for single channel speech separation. However, a typical issue with standard NMF based methods is that they assume the independency of each time frame of speech signal and thus cannot model the temporal continuity of the speech signal. This paper presents a novel algorithm for single-channel speech separation. First, a new model, which combines NMF with k-means clustering method, has been proposed. This model could concurrently describe the spectral structure and the temporal continuity of speech signal. Then the model is used to train factorial conditional random field(FCRF) model, which is used for the separation of mixed speech signal. Experimental results show that the proposed algorithm consistently improves the performance of separation when compared with Active-Set Newton Algorithm(ASNA), a NMF based approach without considering the temporal dynamics of speech signal.
【Key words】 single-channel speech separation; Factorial Conditional Random Field; Non-negative Matrix Factorization; k-means clustering;
- 【会议录名称】 第十三届全国人机语音通讯学术会议(NCMMSC2015)论文集
- 【会议名称】第十三届全国人机语音通讯学术会议(NCMMSC2015)
- 【会议时间】2015-10-25
- 【会议地点】中国天津
- 【分类号】TN912.3
- 【主办单位】中国中文信息学会语音信息专业委员会